"Databricks LLM (APIs)/Model Serving. Thoughts from the Frontline." — @DataEngineeringCentral
Why this is in the vault
A field-level taxonomy of Databricks' LLM-serving options (pay-per-token vs. provisioned throughput) framed entirely around governance, cost-attribution, and observability requirements rather than model quality — directly usable as a checklist the next time RDCO scopes a Databricks-hosted agent build for a client.
⚠️ Sponsorship
Confirmed standing pattern (README, 2026-09-14): "Thanks to Delta for sponsoring this newsletter! I use Delta Lake daily, and I believe it represents the future of Data Engineering." Same daily-user disclosure framing as prior issues. Delta Lake has no editorial role in this issue's actual argument — the piece is about LLM serving tiers and Unity Gateway, Delta is mentioned once more in passing (payload logging destination) — but per the standing-sponsor rule, treat the mention as commercially motivated by default. This is now the 10th+ logged instance; the pattern holds, no break.
The core argument
Daniel Beach lays out a practitioner's decision framework for serving LLM traffic on Databricks, deliberately skipping model-quality debates in favor of the operational question: what do you need besides an inference endpoint?
- Two serving tiers cover most real use cases. Pay-per-token Foundation Model APIs are the "starter pack" — zero setup, OpenAI-spec compatible, but traffic and inference-table logging are pooled at the workspace level (every team's usage mixed together, admin-gated visibility). Provisioned Throughput is the isolation tier — dedicated capacity, your own inference table/catalog/schema/ACLs, batch payload logging, traffic splitting for canary/A-B, and scale-to-zero.
- Unity Gateway is the governance layer that sits in front of either tier, not a third serving option. It centralizes access (UC privileges, ABAC grants, on-behalf-of-user execution for MCP), traffic/cost controls (smart routing, per-user budgets and hard caps, rate limits), guardrail policies (PII redaction, prompt-injection/exfiltration detection, MCP tool-invocation controls returning HTTP 200 with a block reason rather than 4xx), and observability (a unified OpenTelemetry trace table across all gateway services, per-request cost attribution by tag).
- The real architectural decision isn't "which model" — it's who needs isolated logging, who needs traffic splitting, and who's accountable when a Data Scientist blows the token budget. The author frames this as the difference between treating LLM serving as "API + LLM = done" versus treating it as an architectural decision with the same weight as any other production API.
Mapping against Ray Data Co
Direct operational overlap with [[2026-06-19-data-engineering-central-databricks-summit-2026]] (Databricks summit review, same author's beat) and with the AI Gateway pattern RDCO already runs in production for Scribble Works (~/rdco-vault/02-sops/2026-09-07-scribble-works-process-map.md — Cloudflare AI Gateway fronting Claude/Grok calls). The governance vocabulary here — per-tag cost attribution, hard budget caps, blocked-interaction semantics — is a checklist RDCO should be applying to its own gateway rail, not just something to hand a Databricks client: Scribble Works currently logs via Cloudflare AI Gateway but doesn't yet have documented per-feature budget caps or a blocked-interaction contract, which this piece implies is a gap worth closing before the next cost surprise. For phData-adjacent Databricks engagements, the pay-per-token-vs-provisioned-throughput framing (shared inference table + admin-gated logs vs. dedicated table you own) is a concrete talking point for scoping conversations about client agent deployments — "which tier" is now a governance question, not a model-selection question.
Related
[[2026-06-19-data-engineering-central-databricks-summit-2026]] [[2026-04-22-data-engineering-central-most-teams-doing-it-wrong]] [[reference_scribble_works_engine_rail]] [[project_scribble_works_ops_rules]]