06-reference

indy dev dan agentic engineering benchmarks

2026-09-14·reference·source: IndyDevDan (YouTube)·by IndyDevDan
benchmarksagentic-engineeringmodel-selectionharness-engineeringagent-swarms

"Agentic Engineering Benchmarks: How I RANK Astra, Fable 5.1, and Open-Weights" — IndyDevDan

Full transcript: [[2026-09-14-indy-dev-dan-agentic-engineering-benchmarks-transcript]].

Why this is in the vault

IndyDevDan is a tracked author whose harness-engineering framework directly informs RDCO's own agent-fleet design. This episode is a rare explicit statement of how he picks models — a five-benchmark selection method (Terminal-Bench, Apex Agents, Automation Bench, Omniscience, Deep SWE) that replaces reliance on composite indices like Artificial Analysis. The underlying methodology (pick benchmarks that proxy your actual workload, weigh performance/cost/speed as one unit, distrust flat-line "saturated" benchmarks) is a reusable model-selection framework RDCO can apply directly when choosing models for its own agent stack, not just a recap of scores.

Episode summary

Dan argues that composite benchmark indices (Artificial Analysis) compress too much information to be useful for choosing models for a specific workload, and instead proposes his own top-five benchmark stack: Terminal-Bench v4 (pure agentic coding, performance/cost/speed triangle), Apex Agents (non-software knowledge work — investment banking, consulting, corporate law — as a proxy for other domains), Automation Bench (cross-application business automation graded on guardrail-violation-free task completion), Omniscience (hallucination/honesty rate, rewarding "I don't know" over confident wrong answers), and Deep SWE (long-horizon software engineering from short prompts). Across all five he uses GPT-6 Astra as a running control model, repeatedly showing it wins on cost and speed even when tied or slightly behind on raw performance, while Claude Fable 5 / 5.1 trail on cost by 2.5-4x. He closes by naming the throughline across his picks — building toward long-horizon, low-oversight "agentic engineering" (agent swarms, software factories, eventually "dark factories") — and explains why he excluded saturated benchmarks (e.g., a flat-lined long-context retrieval benchmark) and why he advocates a multi-model "model stack" (state-of-the-art / workhorse / lightweight tiers) rather than a single default model.

Key arguments / segments

Notable claims

Guests

N/A — solo creator video (IndyDevDan).

Mapping against Ray Data Co

Directly actionable for RDCO's L5 north star, since RDCO's bets are explicitly downstream of agent capability and model choice. This video is less a benchmark recap than a transferable method: pick a small set of benchmarks that proxy the actual workload (RDCO's is agent-harness/vault-ops/multi-step-write work, closer to Automation Bench's cross-application guardrail-adherence framing and Deep SWE's long-horizon-from-short-prompt framing than to raw SWE-bench coding scores), weigh cost/speed/performance as one unit rather than chasing headline scores, and run a model stack instead of a single default. The Omniscience "zero penalty for I don't know" framing directly validates RDCO's existing sub-agent design pattern (Agent tool returns with explicit "status-only" vs. "DECISION" slots, fresh-eyes critics that can return ITERATE/SCRAP rather than a forced pass) — a hallucinating or over-confident sub-agent silently corrupting a downstream vault write is exactly the failure mode Dan describes for long agent chains. Automation Bench's guardrail-violation framing also maps cleanly onto RDCO's classifier hard-gate (deploy/production-write denial) and the supervise/verify-* critic family, both of which encode "did you break something on your way to completion" as a first-class check rather than a pass/fail-only harness. Practical next step: when RDCO next evaluates swapping or adding a model to its stack (e.g. for station-code-author or dispatch-heavy skills), use this five-benchmark lens rather than the Artificial Analysis composite score alone.

Related