"Do You Actually Need Real-Time Data?" — SeattleDataGuy (Ben Rogojan)
Why this is in the vault
A clean, quotable articulation of a requirements-definition discipline — "real-time" is a word business stakeholders use loosely, and the data engineer's job is to interrogate the actual decision latency before committing to the (expensive) build.
The core argument
- "Real-time" means wildly different things to different stakeholders — sometimes hourly, sometimes daily, relative to whatever their prior baseline was (monthly manual Excel loads feel real-time once you hit hourly).
- Don't ask "do you want real-time data?" (answer is always yes). Ask decision-latency questions instead: "What decision would have been different if you'd had this data 10 minutes earlier?" / "Does someone actively watch this dashboard, or check it at set times?" / "Is there a human in the loop, or an automated process?"
- Genuine sub-second requirements exist (payment fraud is his example) but they share a signature: no human in the loop — a human-reviewed dashboard essentially never needs millisecond freshness.
- Real-time isn't free: pipeline run-frequency compounds fast (24x/day hourly → 288x/day at 5-min → 1,440x/day per-minute), and each step up multiplies the blast radius of backfills and schema changes. He's seen clients pay 10-20x more in compute for real-time setups relative to batch.
- Freshest ≠ most correct: a healthcare claim can be ingested in minutes but isn't "done" — billing, coding, and adjustments still land after the fact, so a real-time pipeline can confidently report a wrong number.
- Freshness need not be uniform: only the specific fields tied to a real decision need to be real-time; core dimensional data can stay batch.
- Closing test: "what decisions become impossible if this data is five minutes old?" — no clear answer means no real-time use case.
⚠️ Sponsorship
Estuary appears in the top-of-article preamble as a co-hosted-event promo ("I am co-hosting an event in Denver with Estuary on October 1st") rather than a direct ad block — this is SDG's disclosed-adviser pattern (per the process-newsletter README gotcha file, appears in ~60% of issues). No explicit self-consulting CTA ("sponsored by me, the Seattle Data Guy") appeared in this issue — checked for both forms per the known gotcha, only the Estuary form is present this time. Estuary is a real-time-ingestion/CDC vendor, so there's a structural (not just financial) reason for SDG to frame real-time-data questions the way he does here — worth reading the "real-time isn't free" section as coming from someone whose adviser relationship is with a company that sells real-time pipelines, which if anything cuts against self-interest (he's arguing customers often shouldn't buy what Estuary sells).
Curation section
"Articles Worth Reading" carries two items, one of each type per SDG's usual ~50/50 split:
- "Using local LLMs for agentic coding" (Alex Ewerlöf) — genuine third-party: different author, personal 3-year local-LLM/hardware retrospective (RTX/M4/ROCm, Llama.cpp/Ollama/LM Studio/Jan), no overlap with SDG's own byline or domain.
- "Backfills - The Necessary Evil of Data Engineering" — self-cross-promo, SDG's own prior piece, already filed at [[2026-02-23-seattle-data-guy-backfills]]. Fittingly on-theme: this issue's "every step up in real-time frequency multiplies your blast radius on backfills/schema changes" point leans directly on that earlier article's argument.
Also namechecked but not linkable: a "Video of the Week" (SDG + Phil Sparks on "Is AI Doom Genuine Fear or Good Marketing?") — mentioned by title only, no URL in the plaintext body, not treated as a curation item.
Mapping against Ray Data Co
Mapping strength: medium. This isn't a direct thesis restatement like [[2026-05-26-seattle-data-guy-ai-consultants-utility-thesis]] (SDG's cleanest external articulation of the agent-deployer wedge), but it's the same underlying discipline applied to a narrower technical question: the real work is translating a vague business ask ("we want real-time") into an actual decision-latency requirement before spending engineering effort — the same "understand the business, the process, the edge cases" move SDG makes in the consultants piece, here at the level of a single infrastructure decision rather than the whole AI-consulting wedge. It's a useful worked example of the underlying skill (interrogate the ask before building) that the agent-deployer thesis assumes deployers actually have, but the note doesn't extend or complicate RDCO's positioning — filing for the pattern-reinforcement, not for new evidence.
Related
- [[2026-05-26-seattle-data-guy-ai-consultants-utility-thesis]] — SDG's direct agent-deployer-thesis restatement; this issue applies the same "interrogate the actual business need before building" discipline to a narrower question
- [[2026-02-23-seattle-data-guy-backfills]] — cross-promoted in this issue's curation section, and thematically linked (backfill/schema-change blast radius scales with pipeline run-frequency, the same tradeoff this issue raises for real-time cadence)
- [[feedback_targeting_system_prioritization_filter]] — the founder's own "can Ray do X" four-layer filter is the same before-you-build interrogation discipline this article argues data engineers owe the business