06-reference/research

agent cost attribution stack

2026-08-29·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
agent-observabilitytoken-attributionopentelemetrylangfuseharness-instrumentation

The Attribution Stack Is Already Installed: Claude Code's Own Telemetry Beats Every Vendor Here

The question

"What tools and frameworks (Langfuse, Helicone, Braintrust, Datadog Agent Observability) best connect LLM token spend to business outcomes — and what does a minimal viable agent-cost-attribution stack look like for a solo-founder COO setup?"

Context: the workload in question is the always-on Ray COO agent on a Mac Mini (cron skills, Discord + iMessage channels, MCP servers, heavy sub-agent fan-out) running on a Claude Max subscription, not metered API billing.

The premise problem, stated plainly

Langfuse, Helicone, Braintrust and Datadog all compute "spend" the same way: capture token counts, multiply by a per-model price table, roll up dollars. That model assumes a meter. A Max subscription has no meter. Claude Code does emit a claude_code.cost.usage metric in USD, but it is computed at public API list price — Anthropic's own cost docs note the figure is "not relevant for billing purposes" for Pro/Max subscribers, and that even API customers on contracted rates will not see their bill in it (code.claude.com/docs/en/costs; corroborated by cloudzero.com, which reports a real case of ~$15k of list-price "cost" against $800 actually paid on Max).

So the question splits cleanly:

What we already know (from the vault)

What the web says (verified against vendor docs, Aug 2026)

Claude Code ships the attribution layer natively, and it is better-shaped for this question than any of the vendors. Per code.claude.com/docs/en/monitoring-usage, setting CLAUDE_CODE_ENABLE_TELEMETRY=1 emits OTLP metrics including claude_code.token.usage (unit: tokens) with attributes type (input/output/cacheRead/cacheCreation), model, query_source (main / subagent / auxiliary), speed, effort, agent.name, skill.name, plugin.name, mcp_server.name, mcp_tool.name, plus session-level claude_code.session.count and claude_code.active_time.total. Exporters: otlp (grpc, http/json, http/protobuf), prometheus, or console. Critical detail: custom skill and agent names are redacted to "third-party" / "custom" unless OTEL_LOG_TOOL_DETAILS=1 is set - and every RDCO skill is custom, so without that flag the whole exercise returns nothing useful.

A large slice of the answer is already on disk, retroactively. Direct inspection of ~/.claude/projects/ (1,594 session JSONL files, ~724MB) confirms every assistant turn already carries message.usage (input / output / cache-creation / cache-read tokens), model, effort, sessionId, cwd, timestamp, and an attributionSkill field. A ~40-line aggregation over the 60 most recent sessions produced a working skill-level breakdown in one pass - top consumers check-board, process-newsletter, deep-research, sync-contacts, open-threads-check, and a model-tier split (Fable > Sonnet > Opus > Haiku by token volume). No vendor, no daemon, no signup.

That local path has one real blind spot, and it is exactly the one the question cares about. No transcript line in the corpus carries isSidechain: true, and Task tool results persist agentId / resolvedModel / status but not the sub-agent's token totals. So transcript-only attribution systematically under-counts sub-agent fan-out - it sees the dispatch, not the cost. OTel's query_source=subagent + agent.name attributes are precisely what closes that gap, and are the strongest single reason to turn telemetry on.

Vendor findings (verified against official docs/pricing, Aug 2026):

Source-bias note: every capability claim above traces to vendor-owned documentation, because no independent documentation exists for proprietary products. Pricing/free-tier figures on marketing product pages (Datadog's 40k-spans tier, Arize AX tiers) are marketing surfaces and were flagged as such. Any "X vs Y" comparison page published by one of these six vendors should be treated as sales collateral, per the vault standard.

Convergences and contradictions

Synthesis for RDCO

Recommendation: do not adopt Langfuse, Helicone, Braintrust, Datadog, Phoenix, or Traceloop. Build the two-hour local thing instead, and hold the vendor decision behind a named trigger.

The reframe that makes this decidable: the question is not "what do the tokens cost," it is "which skills and fan-outs are eating the agent's capacity, and are they earning it." Every dollar-denominated feature in these six products is dead weight against a flat-rate seat, and the one feature that matters - grouping token volume by skill and by sub-agent - Claude Code emits natively and the local transcripts already half-record. Paying a vendor (in setup hours, in a Docker stack on the Mac Mini, in an Elastic-2.0 or /ee license question) to re-render data already sitting in ~/.claude/projects/ is the wrong trade for one person.

Tier 0 - do now, ~2 hours, zero new dependencies. A local aggregator over ~/.claude/projects/*.jsonl that rolls tokens by attributionSkill, model tier, and week, writing a small table into the vault on the existing weekly cadence. It runs retroactively over months of history - no observability vendor can give him that, because they only see traffic from the moment you wire them in. What it tells him that he does not know now: which cron skills dominate consumption (the 60-session sample says check-board, process-newsletter, deep-research, sync-contacts - and open-threads-check fires the most often while producing among the least output, which is a real candidate finding), and whether model-tier routing matches intent (Fable at roughly 8.7M tokens versus Sonnet at 4.9M in that sample is a routing fact worth eyeballing against the delegation policy in feedback_delegation_model_effort_pairing). Pair each skill's token share against its output count - briefs filed, board items closed, contacts synced - and the ratio is the outcome attribution. No dollars appear anywhere in the artifact.

Tier 1 - conditional, ~3-4 hours, only if Tier 0's blind spot bites. CLAUDE_CODE_ENABLE_TELEMETRY=1, OTEL_METRICS_EXPORTER=prometheus (or otlp into a single local collector), and critically OTEL_LOG_TOOL_DETAILS=1 so custom skill and agent names are not redacted to "custom". Leave OTEL_LOG_USER_PROMPTS and OTEL_LOG_RAW_API_BODIES off - prompt bodies contain client and family material and there is no reason to persist them to a metrics store. This is the only path that attributes sub-agent fan-out, which the transcripts cannot see. It needs no vendor: a Prometheus scrape plus Grafana, or the metrics piped to the same aggregation script, is sufficient. Trigger condition: the first time a Tier 0 report shows a skill's consumption is inexplicable from its own transcript, or the first time a fan-out-heavy skill (deep-research, process-newsletter, the brigade stations) is suspected of burning capacity disproportionate to output and the parent transcript cannot settle it.

Tier 2 - the only case where a vendor is the right answer. If RDCO ships a surface that calls a metered API on behalf of someone other than the founder - an API-billed product endpoint, a client-facing agent, a paid Squarely or MAC feature - then per-token dollars become real, unit economics become real, and Langfuse self-hosted (MIT core, OTLP ingest, cost-by-tag) is the pick, chosen on license posture rather than features: it is the only one whose free self-host survives the solo-to-commercial flip without a licensing conversation. Braintrust is closed-core, Phoenix is Elastic-2.0, Traceloop's dashboard is commercial, Datadog and Helicone are structurally wrong for this shape. Note that even then, the Max agent itself should stay outside that instrumentation - mixing an imputed list-price number into a P&L view of real metered spend would produce exactly the wrong picture.

The negative here is well-argued, not lazy: these are good products aimed at a problem RDCO will only have on the other side of a product launch. Buying the instrumentation before the meter exists is the same error the vault already flagged in the data-observability vendors - adding a surface and calling it a measurement.

Why this is in the vault

This closes the open instrumentation question that [[2026-06-01-indy-dev-dan-coding-agent-observability]] left standing ("no trace surface over sub-agent dispatches") and that [[2026-08-04-mostlymetrics-cfo-token-cost-gate]] restated as a missing spend gate, with a decision rather than another restatement: build the Tier 0 transcript aggregator, hold the vendor buy behind the sub-agent-attribution trigger. It also pre-answers the tooling question for the first genuinely API-billed RDCO surface, so that decision does not get re-litigated from scratch under launch pressure.

Open follow-ups

Related

Sources

Vault

Primary / vendor documentation

Secondary (flagged)

Local system inspection (2026-08-29)