"Agentic Engineering Benchmarks: How I RANK Astra, Fable 5.1, and Open-Weights" — IndyDevDan
Full transcript: [[2026-09-14-indy-dev-dan-agentic-engineering-benchmarks-transcript]].
Why this is in the vault
IndyDevDan is a tracked author whose harness-engineering framework directly informs RDCO's own agent-fleet design. This episode is a rare explicit statement of how he picks models — a five-benchmark selection method (Terminal-Bench, Apex Agents, Automation Bench, Omniscience, Deep SWE) that replaces reliance on composite indices like Artificial Analysis. The underlying methodology (pick benchmarks that proxy your actual workload, weigh performance/cost/speed as one unit, distrust flat-line "saturated" benchmarks) is a reusable model-selection framework RDCO can apply directly when choosing models for its own agent stack, not just a recap of scores.
Episode summary
Dan argues that composite benchmark indices (Artificial Analysis) compress too much information to be useful for choosing models for a specific workload, and instead proposes his own top-five benchmark stack: Terminal-Bench v4 (pure agentic coding, performance/cost/speed triangle), Apex Agents (non-software knowledge work — investment banking, consulting, corporate law — as a proxy for other domains), Automation Bench (cross-application business automation graded on guardrail-violation-free task completion), Omniscience (hallucination/honesty rate, rewarding "I don't know" over confident wrong answers), and Deep SWE (long-horizon software engineering from short prompts). Across all five he uses GPT-6 Astra as a running control model, repeatedly showing it wins on cost and speed even when tied or slightly behind on raw performance, while Claude Fable 5 / 5.1 trail on cost by 2.5-4x. He closes by naming the throughline across his picks — building toward long-horizon, low-oversight "agentic engineering" (agent swarms, software factories, eventually "dark factories") — and explains why he excluded saturated benchmarks (e.g., a flat-lined long-context retrieval benchmark) and why he advocates a multi-model "model stack" (state-of-the-art / workhorse / lightweight tiers) rather than a single default model.
Key arguments / segments
- [00:00:00] Opens by naming the problem: the Artificial Analysis index (now v4.3) compresses too many benchmarks into one number, obscuring which specific benchmark matters for a given engineer's work. Proposes the exercise: pick only five benchmarks.

- [00:02:00] Pick #1: Terminal-Bench v4 — pure agentic coding benchmark (60+ tasks across SWE, ML, science, ops, security, hardware, media). Frames model choice as a 3D problem: performance, cost, speed together ("trade-off triangle"). Astra leads on performance; Claude Fable 5.1 trails; sharp drop-off at the 40-42% score band separates state-of-the-art from the pack.

- [00:04:00] Cost breakdown on Terminal-Bench: Astra uses ~2.7x fewer tokens than Claude Fable 5 and ~2.5x fewer than Fable 5.1; introduces his real metric of interest, "useful agent output per hour," not raw score or tokens-per-task.

- [00:07:16] Cost is "where things get really wacky" — Astra roughly 4x cheaper than Fable and 2.4x cheaper than Fable 5.1 on Terminal-Bench despite near-tied headline scores; digs into one sample task (reverse-engineering a UI screenshot into a config.json layout) to show what the benchmark actually measures.

- [00:08:00] Pick #2: Apex Agents — knowledge-work benchmark across investment banking, management consulting, and corporate law, tasks vetted by McKinsey/BCG/Deloitte/Goldman/Morgan Stanley/JPMorgan experts. Used as a proxy for non-software domains RDCO-style businesses actually operate in; flags GLM 5.3 and Kimi K3 surprises (one underperforms, one overperforms expectations).
- [00:11:56] Pick #3: Automation Bench — 600+ cross-application business-automation tasks across finance/HR/marketing/ops/sales/support, graded on completing objectives without triggering guardrail violations — a pass/fail dimension most benchmarks omit entirely. Frames guardrail-adherence as core alignment: "did you break something on your way to completion."

- [00:15:39] Domain-level breakdown on Automation Bench (finance, HR, marketing, etc.) shows model strength varies sharply by application/domain (e.g., Fable weak on finance tasks vs. Astra), reinforcing his "model stack, not one model" thesis — pick different models for different domains rather than defaulting to one.

- [00:20:00] Pick #4: Omniscience — the hallucination/honesty benchmark. Each answer graded correct/incorrect/partial/not-attempted, with zero score penalty for saying "I don't know" — argues this is critical for long-running unsupervised agent chains where one hallucination corrupts every downstream step.

- [00:25:00] Pick #5: Deep SWE (v1.1) — long-horizon software-engineering benchmark using deliberately short, low-detail prompts to test whether an agent can independently expand a "prompt scaled up" into a full plan. Astra wins with ~30K tokens and 29 steps; Gemini 3.8 Flash flagged repeatedly as his favorite cost/performance "workhorse" model.

- [00:29:51] Synthesis begins: names the common thread across all five picks — building toward agentic systems that "run autonomously with no oversight." Explicitly rejects a long-context-retrieval benchmark shown as a flat saturated line ("if I see this pattern, I walk away from the benchmark") as a variance/signal heuristic for benchmark selection generally.

- [00:32:59] States the "AGI you can't pay for is irrelevant" framing — cost-per-task is treated as a first-order filter, not an afterthought, alongside performance and speed; Astra, Gemini 3.8 Flash, GLM 5.3, and Kimi K3 named as current sweet-spot picks.

- [00:34:11] Links this framework directly back to last week's agent-swarm video (OpenAI GPT-6-Astra swarm incident) — argues low deception (Omniscience) and domain-proxy performance (Apex Agents/Automation Bench) are exactly the properties that make swarms and "dark factories" survivable at scale.

- [00:38:00] Closing: names further benchmark gaps he wishes existed (agent-to-agent delegation/handoff quality, small-agent-team coordination, failure recovery) and restates the core takeaway — pick benchmarks that proxy your actual work, not a global index, and run a model stack rather than a single default model.

Notable claims
- Proposes a specific top-5 benchmark stack for agentic engineering: Terminal-Bench v4, Apex Agents, Automation Bench, Omniscience, Deep SWE — explicitly in place of composite indices like Artificial Analysis.
- GPT-6 Astra used as control model across all five benchmarks; reported as ~2.5-4x cheaper than Claude Fable 5/5.1 on Terminal-Bench at near-tied or better performance.
- Automation Bench's key differentiator: grading task completion conditional on zero guardrail violations — a dimension he says most benchmarks omit and considers core to "alignment at a low level."
- Omniscience gives zero score penalty for an agent answering "I don't know" — argued as essential for long-horizon unsupervised agent chains where one hallucination corrupts every downstream step.
- Heuristic for discarding a benchmark: a flat/saturated score line across models ("if I see this pattern, I walk away from the benchmark") signals no discriminating information — offered as a generalizable filter for benchmark selection, not specific to any one benchmark named in the video.
- Restates and extends the agentic-engineering capability ladder from his prior two videos (agent → ADW → software factory → swarm → dark factory → RSI), framing this benchmark stack as measuring the properties (low deception, guardrail adherence, domain generality, long-horizon capability) required to operate at the swarm/dark-factory tiers safely.
Guests
N/A — solo creator video (IndyDevDan).
Mapping against Ray Data Co
Directly actionable for RDCO's L5 north star, since RDCO's bets are explicitly downstream of agent capability and model choice. This video is less a benchmark recap than a transferable method: pick a small set of benchmarks that proxy the actual workload (RDCO's is agent-harness/vault-ops/multi-step-write work, closer to Automation Bench's cross-application guardrail-adherence framing and Deep SWE's long-horizon-from-short-prompt framing than to raw SWE-bench coding scores), weigh cost/speed/performance as one unit rather than chasing headline scores, and run a model stack instead of a single default. The Omniscience "zero penalty for I don't know" framing directly validates RDCO's existing sub-agent design pattern (Agent tool returns with explicit "status-only" vs. "DECISION" slots, fresh-eyes critics that can return ITERATE/SCRAP rather than a forced pass) — a hallucinating or over-confident sub-agent silently corrupting a downstream vault write is exactly the failure mode Dan describes for long agent chains. Automation Bench's guardrail-violation framing also maps cleanly onto RDCO's classifier hard-gate (deploy/production-write denial) and the supervise/verify-* critic family, both of which encode "did you break something on your way to completion" as a first-class check rather than a pass/fail-only harness. Practical next step: when RDCO next evaluates swapping or adding a model to its stack (e.g. for station-code-author or dispatch-heavy skills), use this five-benchmark lens rather than the Artificial Analysis composite score alone.
Related
- [[2026-09-07-indy-dev-dan-agent-swarms-gpt6-astra]] — prior week's video this episode explicitly references and builds on; both videos treat GPT-6 Astra as the reigning control model and continue the same agentic-engineering capability-ladder narrative
- [[2026-08-31-indy-dev-dan-agentic-operating-level]] — originates the agent → ADW → software factory → dark factory → RSI ladder this video's synthesis section extends
- [[2026-08-24-indy-dev-dan-intelligence-explosion-harness-engineering]] — earlier piece on harness-engineering discipline underlying the "pick benchmarks that proxy your workload" methodology here