Getting Started With Open Models — Kai Zau
Why this is in the vault
A practical decision framework — "commodity intelligence" — for when a task needs a frontier model versus when a cheaper or open-weight model is good enough, from a programmer who runs millions of tokens a day on a self-hosted Mac Studio.
The core argument
Kai Zau (intro by Every editor Kate Lee) argues most AI work isn't frontier-grade work: it's routine, repeatable tasks for which a $20/month-tier frontier model is overkill firepower, "a lot of firepower for summarizing your meeting notes." His framework for picking a tier asks: how frequent is the task, how clear are its inputs/outputs, how verifiable is the result, and is the frontier premium economically justified? Poorly-defined, one-off, high-consequence work needing human-grade judgment stays on frontier models; clear, repeatable, easy-to-verify work moves to open or cheaper models. He sketches a model-tier ladder (Extra Large down to Tiny, specifics paywalled) and notes that smaller models become reliable for harder tasks given good instructions and a verification mechanism.
Two points stand out from what rendered before the paywall: (1) the most capable open models currently come from Chinese labs (GLM, etc.), but running them locally or via a non-Chinese host doesn't mean sending data to China — a conflation he explicitly separates; (2) running a mixed stack (frontier + hosted open + self-hosted open) buys redundancy — "when an API goes down, you have a backup" — at the cost of more operational overhead than a single proprietary stack. The adoption-trend framing up top cites Vercel's AI Gateway (56% of August tokens ran through open models), AT&T (40% of AI use on open models, trending to 60%), and Apple's new on-device-AI-forward Mac Studio/Mac Mini as evidence this isn't a fringe practice.
The bulk of the piece — specific hosting providers, tool/coding-agent integration steps, and a hardware-investment decision prompt — sits behind Every's paywall and isn't reproduced here.
Mapping against Ray Data Co
Concrete connection: RDCO's own agent fleet already does informal task-tiering inside the Anthropic stack — the fanout agent type (Sonnet 5, low effort, for "cheap mechanical per-item transforms") versus general-purpose/Opus for judgment work is exactly the frequency/clarity/verifiability split Zau formalizes. What RDCO hasn't done is ask his harder question: of the bounded, high-frequency, easily-verified work already routed to fanout (one-subagent-per-article newsletter processing, mechanical field extraction, classification), how much could run on an open-weight model off Anthropic entirely, cutting the per-call cost that feedback_api_cost_budget_controlled currently treats as invisible-under-Max-sub rather than optimized? This doesn't change anything about RDCO's Anthropic-exclusive posture today — it's a question RDCO's own task-tiering logic implies but has never evaluated, and is worth a deliberate "no, not yet, because X" rather than default inertia.
⚠️ Sponsorship
No third-party sponsor in this issue's body or research disclosures. The footer carries Every's standard house self-promotion: an "Every All Access" membership + Builder Pack upsell ($9,000+ in credits) and a bundled-software pitch for Every's own products (Sparkle, Cora, Spiral, Monologue). Flagged per RDCO convention (sponsor_entity: house-promo) even though it's self-promotion, not paid third-party advertising — it's still a commercial incentive embedded in editorial content.
Related
- [[2026-08-18-technically-open-weight-models]]
- [[2026-09-01-technically-open-weight-models-part2]]
- [[2026-07-19-alphasignal-inkling-pragmatic-open-weights-engineering]]
- [[2026-09-17-every-burn-more-tokens]]
- [[feedback_api_cost_budget_controlled]]
- [[feedback_delegation_model_effort_pairing]]