06-reference

data engineering weekly agentic ml llm judge

2026-09-07·reference·source: Data Engineering Weekly·by Ananth Packkildurai
data-engineeringagentic-aillm-as-judgeml-platformcuration

Data Engineering Weekly #286 — Agentic ML Modeling & LLM-as-a-Judge

Source: https://www.dataengineeringweekly.com/p/data-engineering-weekly-286

Why this is in the vault

Two curated pieces (Instacart's Griffin ML platform, Netflix's LLM-as-a-Judge lifecycle) both describe production patterns for putting an AI agent in the evaluation/generation loop under human supervision — directly on the through-line this vault has been tracking across recent DEW issues (agent ontology, agent coordination, agent-ready data architecture).

Curation section

Two items would have been deep-fetched (cap of 2), both blocked by 403:

1. Instacart — Agentic Machine Learning Modeling at Instacart. Blurb-level detail only (WebFetch 403'd on tech.instacart.com): agents in Griffin autonomously generate, code, and evaluate ML hypotheses under human supervision, going beyond hyperparameter tuning to test different model families and compound architectural improvements from real-time eval feedback.

2. Netflix — The Lifecycle of LLM-as-a-Judge. Blurb-level detail only (WebFetch 403'd on netflixtechblog.medium.com): four-phase lifecycle (Birth, Training, Deployment, Monitoring) for a human-anchored judge system, with Reasoning-Aligned Rubric Tuning (RART) specifically targeting "right label, wrong reason" errors by tuning judge prompts against written human rationales rather than label-agreement alone.

Mapping against Ray Data Co

The Netflix RART framing — catching "right label, wrong reason" errors by grounding judge-prompt tuning in written human rationales, not just label agreement — is a sharper eval-quality bar than anything currently named in RDCO's own harness-engineering vocabulary (which so far talks about pass/fail gates and critic verdicts, not judge-reasoning alignment). Worth testing whether any of the fresh-eyes critic subagents (verify-vault-write, verify-strategic-output) could adopt a rubric-tuning step instead of a static rubric, the next time one of those critics is revised. Instacart's Griffin pattern — agents doing ML hypothesis generation/testing under human supervision rather than parameter search — is a concrete "agent does the work, human gates it" example RDCO can cite in client conversations about what agentic ML actually looks like in production versus the buzzword version.

Related

⚠️ Sponsorship

Two self-interest markers in this issue: (1) the lead item is Data Engineering Weekly's own eBook ("Data Platform Fundamentals") — house promotion, not a paid third party, and identical to the lead slot in issue #285. (2) A generic "AI Modernization Guide" mid-issue is explicitly labeled "Sponsored:" with no attributed author or publication — same slot and generic format as the sponsored guide in issue #285, suggesting a recurring/standing ad unit rather than a fresh placement. Neither sponsor/house item was deep-fetched or used in the mapping above.