The Attribution Stack Is Already Installed: Claude Code's Own Telemetry Beats Every Vendor Here
The question
"What tools and frameworks (Langfuse, Helicone, Braintrust, Datadog Agent Observability) best connect LLM token spend to business outcomes — and what does a minimal viable agent-cost-attribution stack look like for a solo-founder COO setup?"
Context: the workload in question is the always-on Ray COO agent on a Mac Mini (cron skills, Discord + iMessage channels, MCP servers, heavy sub-agent fan-out) running on a Claude Max subscription, not metered API billing.
The premise problem, stated plainly
Langfuse, Helicone, Braintrust and Datadog all compute "spend" the same way: capture token counts, multiply by a per-model price table, roll up dollars. That model assumes a meter. A Max subscription has no meter. Claude Code does emit a claude_code.cost.usage metric in USD, but it is computed at public API list price — Anthropic's own cost docs note the figure is "not relevant for billing purposes" for Pro/Max subscribers, and that even API customers on contracted rates will not see their bill in it (code.claude.com/docs/en/costs; corroborated by cloudzero.com, which reports a real case of ~$15k of list-price "cost" against $800 actually paid on Max).
So the question splits cleanly:
- Not answerable, and shouldn't be attempted: dollar attribution for the Max agent. Any dollar number a vendor tool renders here is a fictional API-equivalent, and reporting it violates the standing no-dollar-figure rule (
feedback_api_cost_budget_controlled). - Answerable, and genuinely useful: effort attribution. Tokens, turns, wall-clock and model tier, grouped by which skill / sub-agent / MCP server / cron job consumed them. That is a real, unitless-but-comparable measure of where the agent's capacity goes, and it is what "connect spend to outcomes" actually means when the spend is fixed and the scarce resource is throughput and context.
- Separately answerable, small today: RDCO's genuinely API-billed surfaces (ElevenLabs TTS, image/video render calls, any future product surface calling the Anthropic API directly). Those have real meters and are the only place per-token dollar attribution is honest. Today they are a rounding error against the subscription, and none of them route through Claude Code.
What we already know (from the vault)
- The gap is already named. [[2026-06-01-indy-dev-dan-coding-agent-observability]] concluded RDCO has "no trace surface over what those sub-agents cost, how many turns they took, or what their rendered system prompt actually contained," and recommended a lightweight per-dispatch trace log rather than a dashboard: "adopt the diagnostic discipline, not the dashboard, until dispatch volume justifies it."
- The economic frame is already set. [[2026-08-04-mostlymetrics-cfo-token-cost-gate]] gives the COGS/opex split (production work vs internal research crons) and the "drive to true per-unit numbers - cost per ticket resolved, per document reviewed" prescription, and explicitly flags that RDCO has zero spend checkpoint because the Max subscription hides the meter.
- [[2026-06-23-every-token-tightening]] frames the same thing from the demand side: enterprises are moving to ROI-gated token allocation, which "is a forcing function toward better observability... the agent needs logging that connects sessions to business outputs."
- The vendor-motion skepticism is precedented. [[2026-07-11-monte-carlo-ai-observability-vs-mac]] found the 2026 AI-observability wave to be mostly "movement on the y-axis of the market (more layers covered) without movement on the x-axis (how tests get designed)" - anomaly-detection mechanics ported to LLM signals. That prior applies directly here: these tools add a surface, not a new kind of measurement.
- Langfuse is already a known name in the vault, but as an evals source, not an observability purchase ([[2026-05-20-lotte-verheyden-evals-explained-langfuse-academy]], which fed the two-gate verdict architecture in [[2026-05-20-verify-stack-two-gate-pass-fail-architecture]]).
- Subscription posture is settled and load-bearing: [[2026-07-04-anthropic-max-plan-tos-productized-agentic-use]] governs how far the Max substrate can be pushed. Nothing in this brief proposes changing it.
What the web says (verified against vendor docs, Aug 2026)
Claude Code ships the attribution layer natively, and it is better-shaped for this question than any of the vendors. Per code.claude.com/docs/en/monitoring-usage, setting CLAUDE_CODE_ENABLE_TELEMETRY=1 emits OTLP metrics including claude_code.token.usage (unit: tokens) with attributes type (input/output/cacheRead/cacheCreation), model, query_source (main / subagent / auxiliary), speed, effort, agent.name, skill.name, plugin.name, mcp_server.name, mcp_tool.name, plus session-level claude_code.session.count and claude_code.active_time.total. Exporters: otlp (grpc, http/json, http/protobuf), prometheus, or console. Critical detail: custom skill and agent names are redacted to "third-party" / "custom" unless OTEL_LOG_TOOL_DETAILS=1 is set - and every RDCO skill is custom, so without that flag the whole exercise returns nothing useful.
A large slice of the answer is already on disk, retroactively. Direct inspection of ~/.claude/projects/ (1,594 session JSONL files, ~724MB) confirms every assistant turn already carries message.usage (input / output / cache-creation / cache-read tokens), model, effort, sessionId, cwd, timestamp, and an attributionSkill field. A ~40-line aggregation over the 60 most recent sessions produced a working skill-level breakdown in one pass - top consumers check-board, process-newsletter, deep-research, sync-contacts, open-threads-check, and a model-tier split (Fable > Sonnet > Opus > Haiku by token volume). No vendor, no daemon, no signup.
That local path has one real blind spot, and it is exactly the one the question cares about. No transcript line in the corpus carries isSidechain: true, and Task tool results persist agentId / resolvedModel / status but not the sub-agent's token totals. So transcript-only attribution systematically under-counts sub-agent fan-out - it sees the dispatch, not the cost. OTel's query_source=subagent + agent.name attributes are precisely what closes that gap, and are the strongest single reason to turn telemetry on.
Vendor findings (verified against official docs/pricing, Aug 2026):
- Langfuse - MIT-licensed core, self-hostable free with no license key (
git clone+docker compose, ~10-15 min); generic OTLP-HTTP ingest at/api/public/otel(traces only, no gRPC). Supports tags/metadata/sessionIdand a cost-by-tag view, and allowscost_detailsoverride at ingestion, so a $0 / flat-rate convention is expressible. The/eefolder (RBAC, SSO/SCIM, audit logs, retention policies, data masking) is separately licensed and commercial regardless of headcount. Its documented Claude Code integration is a Stop hook that parses transcripts and pushes via the Langfuse SDK - not OTLP, i.e. it re-implements the local-transcript path with a server in front of it. Cloud free tier: 50k units/mo, 2 seats, 30-day retention. - Helicone - Apache-2.0, genuinely free self-host (
docker run helicone-all-in-one), but the architecture is a proxy: you pointbaseURLat their gateway. There is no documented generic OTLP endpoint, and Claude Code on a subscription does not route through a user-controlled base URL. Structurally inapplicable to this workload. No documented custom/zero-pricing mechanism. - Braintrust - real generic OTLP endpoint (
https://api.braintrust.dev/otel/v1/traces) and per-spanmetrics.estimated_costoverride, but the core platform is closed-source SaaS, not self-hostable below Enterprise BYOC. Free Starter tier is 1GB / 10k scores / 14-day retention, owner-only. Braintrust is an evals product with tracing attached; RDCO's eval surface is the verify-* critic family, which is already built and does not need it. - Datadog LLM Observability - SaaS-only, no self-host, free tier ~40k spans/mo with 15-day retention (figure appears on the marketing product page, not core docs). OTLP ingest works only if spans follow OTel GenAI semconv/OpenInference with a
dd-otlp-source=llmobsheader. No native Claude Code trace integration; the closest artifacts are "Lapdog" (a separate local dev tool) and Claude Code Skills for querying data you already instrumented. Cost model is a per-token catalog with no flat-rate mode. Massive overkill for one person and one machine. - Arize Phoenix - the closest technical fit after native OTel: true OTLP-HTTP (
:6006/v1/traces) and gRPC (:4317),pip install arize-phoenix && phoenix serve, per-model price table overridable in settings. License is Elastic License 2.0 - fine for solo/internal self-host, but it forbids offering the software as a hosted service to third parties and forbids circumventing license-key gating, so it is the one entry in the set with real teeth against the "flips to commercial on external distribution" rule. No cost-by-tag view found. - OpenLLMetry / Traceloop - SDK is Apache-2.0 and backend-agnostic (thin OTLP wrapper), but it instruments your code calling an LLM SDK, which is not the shape of Claude Code. Traceloop's own dashboard is closed commercial even self-hosted, needs Kubernetes + ClickHouse + Kafka + Postgres, and the free cloud tier has 24-hour retention. Per-model dollar pricing was still an open GitHub issue.
Source-bias note: every capability claim above traces to vendor-owned documentation, because no independent documentation exists for proprietary products. Pricing/free-tier figures on marketing product pages (Datadog's 40k-spans tier, Arize AX tiers) are marketing surfaces and were flagged as such. Any "X vs Y" comparison page published by one of these six vendors should be treated as sales collateral, per the vault standard.
Convergences and contradictions
- Convergence: the vault's own prior (instrument the diagnostic, not the dashboard - [[2026-06-01-indy-dev-dan-coding-agent-observability]]) and the vendor survey land in the same place from opposite directions. The tools are competent; they are just solving a metered-API problem RDCO does not have.
- Contradiction worth naming: [[2026-08-04-mostlymetrics-cfo-token-cost-gate]] argues for "building a gate" on token spend. That argument is sound for a company with a bill and wrong for RDCO right now. On a flat-rate seat the binding constraint is not dollars, it is context budget and founder attention. A gate that pauses work on cost would directly violate
feedback_api_cost_budget_controlledand would optimize the wrong scarce resource. - Mild contradiction inside Langfuse's own material: it markets OTLP-native ingest, but its published Claude Code path is a transcript-parsing Stop hook. That is a tell that Claude Code telemetry is not yet a first-class ingestion story for any of these vendors - which is an argument for waiting, not for picking one.
Synthesis for RDCO
Recommendation: do not adopt Langfuse, Helicone, Braintrust, Datadog, Phoenix, or Traceloop. Build the two-hour local thing instead, and hold the vendor decision behind a named trigger.
The reframe that makes this decidable: the question is not "what do the tokens cost," it is "which skills and fan-outs are eating the agent's capacity, and are they earning it." Every dollar-denominated feature in these six products is dead weight against a flat-rate seat, and the one feature that matters - grouping token volume by skill and by sub-agent - Claude Code emits natively and the local transcripts already half-record. Paying a vendor (in setup hours, in a Docker stack on the Mac Mini, in an Elastic-2.0 or /ee license question) to re-render data already sitting in ~/.claude/projects/ is the wrong trade for one person.
Tier 0 - do now, ~2 hours, zero new dependencies. A local aggregator over ~/.claude/projects/*.jsonl that rolls tokens by attributionSkill, model tier, and week, writing a small table into the vault on the existing weekly cadence. It runs retroactively over months of history - no observability vendor can give him that, because they only see traffic from the moment you wire them in. What it tells him that he does not know now: which cron skills dominate consumption (the 60-session sample says check-board, process-newsletter, deep-research, sync-contacts - and open-threads-check fires the most often while producing among the least output, which is a real candidate finding), and whether model-tier routing matches intent (Fable at roughly 8.7M tokens versus Sonnet at 4.9M in that sample is a routing fact worth eyeballing against the delegation policy in feedback_delegation_model_effort_pairing). Pair each skill's token share against its output count - briefs filed, board items closed, contacts synced - and the ratio is the outcome attribution. No dollars appear anywhere in the artifact.
Tier 1 - conditional, ~3-4 hours, only if Tier 0's blind spot bites. CLAUDE_CODE_ENABLE_TELEMETRY=1, OTEL_METRICS_EXPORTER=prometheus (or otlp into a single local collector), and critically OTEL_LOG_TOOL_DETAILS=1 so custom skill and agent names are not redacted to "custom". Leave OTEL_LOG_USER_PROMPTS and OTEL_LOG_RAW_API_BODIES off - prompt bodies contain client and family material and there is no reason to persist them to a metrics store. This is the only path that attributes sub-agent fan-out, which the transcripts cannot see. It needs no vendor: a Prometheus scrape plus Grafana, or the metrics piped to the same aggregation script, is sufficient. Trigger condition: the first time a Tier 0 report shows a skill's consumption is inexplicable from its own transcript, or the first time a fan-out-heavy skill (deep-research, process-newsletter, the brigade stations) is suspected of burning capacity disproportionate to output and the parent transcript cannot settle it.
Tier 2 - the only case where a vendor is the right answer. If RDCO ships a surface that calls a metered API on behalf of someone other than the founder - an API-billed product endpoint, a client-facing agent, a paid Squarely or MAC feature - then per-token dollars become real, unit economics become real, and Langfuse self-hosted (MIT core, OTLP ingest, cost-by-tag) is the pick, chosen on license posture rather than features: it is the only one whose free self-host survives the solo-to-commercial flip without a licensing conversation. Braintrust is closed-core, Phoenix is Elastic-2.0, Traceloop's dashboard is commercial, Datadog and Helicone are structurally wrong for this shape. Note that even then, the Max agent itself should stay outside that instrumentation - mixing an imputed list-price number into a P&L view of real metered spend would produce exactly the wrong picture.
The negative here is well-argued, not lazy: these are good products aimed at a problem RDCO will only have on the other side of a product launch. Buying the instrumentation before the meter exists is the same error the vault already flagged in the data-observability vendors - adding a surface and calling it a measurement.
Why this is in the vault
This closes the open instrumentation question that [[2026-06-01-indy-dev-dan-coding-agent-observability]] left standing ("no trace surface over sub-agent dispatches") and that [[2026-08-04-mostlymetrics-cfo-token-cost-gate]] restated as a missing spend gate, with a decision rather than another restatement: build the Tier 0 transcript aggregator, hold the vendor buy behind the sub-agent-attribution trigger. It also pre-answers the tooling question for the first genuinely API-billed RDCO surface, so that decision does not get re-litigated from scratch under launch pressure.
Open follow-ups
- Does
claude_code.token.usageattribute tokens consumed by a background sub-agent (therun_in_backgroundpath) toquery_source=subagentin real time, or only at join, and does the parent session id survive on those datapoints? The docs list the attribute but not the emission timing for detached agents; this determines whether Tier 1 can actually see async fan-out. - Is there a defensible non-dollar unit for "outcome per unit of agent effort" that survives comparison across dissimilar skills (a research brief vs a board item closed vs a contact synced)? Without one, the Tier 0 ratio is only comparable within a skill over time, not across skills.
- How stable are the OpenTelemetry GenAI semantic conventions as of late 2026, and has
claude_code.*converged toward them or stayed proprietary? This governs whether any Tier 1 metrics store is portable or becomes a one-vendor lock-in by another name.
Related
- [[2026-06-01-indy-dev-dan-coding-agent-observability]] - the vault's prior on agent observability; recommended diagnostic discipline over dashboards, which this brief operationalizes
- [[2026-08-04-mostlymetrics-cfo-token-cost-gate]] - the COGS/opex split and "build a gate" argument this brief accepts in principle and rejects for the current flat-rate case
- [[2026-06-23-every-token-tightening]] - ROI-gated token allocation as the forcing function toward outcome-linked logging
- [[2026-07-11-monte-carlo-ai-observability-vs-mac]] - the y-axis/x-axis vendor-motion prior applied here to LLM observability tooling
- [[2026-07-12-data-observability-cross-system-reconciliation-vs-mac]] - companion brief on where observability vendors stop short of reconciliation
- [[2026-07-04-anthropic-max-plan-tos-productized-agentic-use]] - the subscription-substrate posture that makes the no-meter premise binding
- [[2026-05-20-lotte-verheyden-evals-explained-langfuse-academy]] - Langfuse's role in the vault to date is evals pedagogy, not an observability purchase
- [[2026-05-20-verify-stack-two-gate-pass-fail-architecture]] - the existing eval/critic surface, which is why Braintrust's eval pitch is redundant here
- [[2026-04-15-thariq-claude-code-session-management-1m-context]] - context rot; the reason context budget, not dollars, is the scarce resource being attributed
- [[2026-07-21-alphasignal-llm-fingerprinting-cursor-agent-costs]] - Cursor's planner/worker split as the archetype of a routing decision that only measurement can validate
Sources
Vault
06-reference/2026-06-01-indy-dev-dan-coding-agent-observability.md06-reference/2026-08-04-mostlymetrics-cfo-token-cost-gate.md06-reference/2026-06-23-every-token-tightening.md06-reference/research/2026-07-11-monte-carlo-ai-observability-vs-mac.md06-reference/research/2026-07-12-data-observability-cross-system-reconciliation-vs-mac.md06-reference/research/2026-07-04-anthropic-max-plan-tos-productized-agentic-use.md06-reference/2026-05-20-lotte-verheyden-evals-explained-langfuse-academy.md02-sops/2026-05-20-verify-stack-two-gate-pass-fail-architecture.md06-reference/2026-07-21-alphasignal-llm-fingerprinting-cursor-agent-costs.md06-reference/2026-04-15-thariq-claude-code-session-management-1m-context.md
Primary / vendor documentation
- Claude Code monitoring & OpenTelemetry reference - https://code.claude.com/docs/en/monitoring-usage
- Claude Code cost management (list-price caveat) - https://code.claude.com/docs/en/costs
- Claude Agent SDK observability - https://code.claude.com/docs/en/agent-sdk/observability
- Langfuse OpenTelemetry ingest - https://langfuse.com/integrations/native/opentelemetry
- Langfuse Claude Code integration (Stop-hook path) - https://langfuse.com/integrations/developer-tools/claude-code
- Langfuse licensing / self-hosting - https://langfuse.com/self-hosting/license-key and https://langfuse.com/pricing
- Helicone self-host + gateway docs - https://docs.helicone.ai
- Braintrust OpenTelemetry integration - https://www.braintrust.dev/docs/integrations/sdk-integrations/opentelemetry
- Datadog LLM Observability (OTLP ingest, GenAI semconv) - https://docs.datadoghq.com/llm_observability/
- Arize Phoenix docs / Elastic License 2.0 - https://arize.com/docs/phoenix
- OpenLLMetry / Traceloop - https://www.traceloop.com/docs/openllmetry
Secondary (flagged)
- CloudZero, Claude Code pricing 2026 (list-price vs actual-bill gap; third-party blog, corroborating) - https://www.cloudzero.com/blog/claude-code-pricing/
Local system inspection (2026-08-29)
~/.claude/projects/- 1,594 session JSONL files, ~724MB; per-turnmessage.usage,model,effort,attributionSkill,attributionMcpServer; zeroisSidechain: truelines;Tasktool results carryagentId/resolvedModel/statusbut no sub-agent token totals~/.claude/settings.json- noCLAUDE_CODE_ENABLE_TELEMETRYorOTEL_*variables set; telemetry currently off