Why this is in the vault
Issue #288 curates dbt Summit takeaways plus eight engineering write-ups (Airbnb, Canva, Booking.com, Orb, Bolt, an infra-cost tip, and a "big data or not" essay) — the sharpest item is a proposal for governing AI-agent data access via version-controlled semantic context, a direct hit on RDCO's own agent-governance framing.
Mapping against Ray Data Co
The load-bearing item is Joanna He's "Beyond the Semantic Layer: Engineering the Agentic Data Stack": agents can't reliably work from raw tables and undocumented business rules, so the proposed fix keeps metrics, relationships, permissions, and business context as version-controlled files that agents read through a controlled gateway before querying — the same shape as RDCO's own vault + knowledge-graph discipline (qmd + graph-ingest) enforcing that an agent's working context is explicit and versioned rather than tribal knowledge re-guessed each session. This is the second consecutive DEW issue converging on "agent reads the mess, writes/reads the governed version" (see #287's BlaBlaCar item) — worth treating as a recognized pattern when RDCO pitches agent-deployment work to a technical buyer, not a one-off. Secondary relevance: the editor's own skepticism of the DAVE-stack article ("do we really have a big data problem?") is a useful gut-check against over-architecting RDCO's own tooling before the actual data volume justifies it — a discipline note more than a technical one.
Curation section
- Mark Rittman — dbt Summit 2026 recap: dbt Labs/Fivetran/agentic-data-stack announcements; editor flags dbt's embedded "charts" feature as worth trying, and notes the oddity of seeing managed lakehouse, Fivetran, and open data infra all pitched on one roadmap slide. https://blog.rittmananalytics.com/dbt-summit-2026-whats-new-and-what-s-coming-from-dbt-labs-fivetran-and-the-agentic-data-stack-52bbef9334a9
- Tim Poterba — The DAVE stack (Datafusion/Arrow/Vortex/Embedded): editor is skeptical of the framing itself, redirecting to the underlying question — do you actually have a big-data problem, or just an incremental-processing design problem — using a genetic-analytics domain system (Phoebe) as the counter-example. https://sequenceandsilicon.substack.com/p/the-dave-stack-the-age-of-domain
- Joanna He — Beyond the Semantic Layer: Engineering the Agentic Data Stack: version-controlled metrics/relationships/permissions/context, read by agents through a controlled gateway, to reduce wrong answers and keep access governable. See mapping above. https://medium.com/@Joannahe/beyond-the-semantic-layer-engineering-the-agentic-data-stack-98fbaee9ad10
- Airbnb — Chronon-powered real-time guest journey: search recommendations now react to in-session browsing/search events instead of waiting for a daily batch refresh, keeping the model itself off the fast search path. https://medium.com/airbnb-engineering/the-guest-journey-updated-in-real-time-extending-airbnbs-sequence-recommender-with-chronon-8f1582578553
- Canva — Worker backpressure: queue workers self-throttle concurrency as a downstream dependency's error rate rises, then ramp back up as it recovers, protecting a failing dependency instead of flooding it. https://www.canva.dev/blog/engineering/worker-backpressure-part-1-how-we-taught-our-queue-workers-to-slow-down/
- Booking.com — Automated Kafka capacity testing: replaces manual failure drills with an automated test that shifts extra partitions to one consumer and auto-reverts on failure, producing repeatable evidence of consumer headroom. https://medium.com/booking-com-development/how-we-built-automated-capacity-testing-for-kafka-consumers-1853623bce78
- Orb — Debugging billing data with DuckDB: queries rejected/duplicate billing-usage files in place in S3 via DuckDB rather than copying into a separate analytics system, speeding root-cause of billing mismatches. https://www.withorb.com/blog/debugging-large-datasets-with-duckdb
- Ian Binder — Stop paying NAT Gateway prices for S3 traffic: a free S3 gateway endpoint routes private-subnet traffic to S3 internally, but only if attached to every relevant route table — an easy-to-miss cost fix. https://medium.com/@ianbinder/stop-paying-nat-gateway-prices-for-s3-traffic-3944ccd64bf8
- Bolt Labs — Migrating 10M monthly Looker queries from Presto to Databricks: built a tool to translate both SQL and Looker's Liquid templates, validated dashboards in parallel on both engines, and kept every step reversible during the cutover. https://medium.com/bolt-labs/migrating-10-million-monthly-queries-how-bolt-moved-looker-from-presto-to-databricks-f427666a0de6
No deep-fetches this issue — each blurb (Joanna He's included) already names the concrete mechanism and result, enough to assess relevance without following the link.
⚠️ Sponsorship
One paid third party plus one house self-promo:
- Unnamed vendor, "AI Modernization Guide" — standalone "Sponsored" block pitching a free guide on future-proofing legacy data infra for AI; the sponsoring company's name isn't in the plaintext body (likely a logo in the HTML render), consistent with the same unresolved-sponsor gap noted in issue #287.
- House self-promo: the issue opens with the author's own eBook ("Data Platform Fundamentals") pitched as a lead item, not flagged as an ad. Not a paid third party, but shapes the issue's framing; disclosed for completeness, not counted as bias on the curated third-party items. No cross-promotion of sister publications or the author's own bylines detected among the curated links — all nine linked domains (rittmananalytics.com, sequenceandsilicon.substack.com, medium.com/@Joannahe, airbnb-engineering, canva.dev, booking-com-development, withorb.com, medium.com/@ianbinder, bolt-labs) are independent third parties with no overlap with dataengineeringweekly.com.
Related
- [[2026-09-14-dataengineeringweekly-ai-ready-data-embeddings-platforms]]
- [[2026-08-31-dataengineeringweekly-285-agent-ready-data-architecture]]
- [[2026-09-07-data-engineering-weekly-agentic-ml-llm-judge]]