Assessment note: [[2026-08-24-indy-dev-dan-intelligence-explosion-harness-engineering]].
What's up engineers? Indie Deb Dan here. Kimmy K3 was released mid July. Deepseek V4 Flash late July early August. Quinn 3.8 early August. Muse Glimmer. Neotron 3.5. Grock 4.6. Deepseek V4 Pro. Gemini 3.7 Flash. Quinn 3.8 27 billion and GLM 5.3. We have just experienced an intelligence explosion. We have more than five model releases within 5 days across all tiers of models. The big AI labs are clearly panicked. Fable is a permanent part of the clawed subscriptions now and OpenAI recently slashed prices for Terra Luna and now they're testing 50% cuts on GPT 5.6 sold on Open Router. The LLM pricing wars are in full effect. The intelligence explosion emphasizes two key questions for us engineers. What's the best way to use these models together to outperform them individually? and how can we build for an age of rapid change
[00:01:01] so we can leverage the models with the best performance, speed, cost trade-offs for our valuable engineering work. I've been engineering for over 15 years and in technological revolutions where rapid change is the norm, one engineering principle stands far above the rest. The most flexible system wins. Today I want to share my thoughts on the intelligence explosion and share a V2 of one of my favorite custom PI coding agents, the fusion harness. We're going to talk about why the fusion harness is so valuable and then upscale all of it to talk about inloop agent coding and outloop agentic coding. There are three commands I want to share with you to give you concrete ideas on how to use your flexible agent harness more effectively.
[00:02:00] Before I showcase these, we need to address the intelligence explosion. I maintain a running model stack of the best models — new releases across the lightweight tier (Muse Glimmer, Iron 3.5 Lightning from Nvidia, Quinn 3.8 27B), and moving up into the daily-workhorse tier.
[00:03:00] GLM 5.3 (waiting for a US-based provider), Deepseek V4 Flash (very cheap, very strong), Deepseek V4 Pro, Gemini 3.7 Flash (his favorite model out of the whole explosion), plus Quinn 3.8, Grock 4.6, and Kimmy K3 entering S-tier as open-weights models.
[00:04:00] These frontier open-weight models are basically impossible to self-host without heavy capital, but providers like Fireworks and Open Router make them available at a fraction of state-of-the-art (Fable 5, Opus 5, GPT 5.6) pricing. Demo setup: Fable 5, Gemini 3.7 Flash, and Deepseek V4 Pro run side by side in the "fusion harness" against three tasks, starting with understanding DuckDB's new V2 preview release.
[00:05:00]
First command: /FH opinion — asks every model in the stack the same question (what's the most important thing to focus on in the DuckDB V2 release) and returns each model's answer plus cost/speed. Model aliases (not real names) are used deliberately — revealing a model's real name to peer models causes it to start competing/behaving oddly.
[00:06:01] Three things to weigh when using models together: performance, speed, cost. Gemini 3.7 Flash is described as the fastest model in existence right now; Deepseek V4 Pro is mid-speed; Fable 5 is slowest, most powerful, and an order of magnitude more expensive.
[00:07:00] All three models converge on recommending DuckDB's new "variant" feature. The pitch: instead of running one model for feedback, run N models and understand what one unit of result costs, because that cost scales with your agent usage.
[00:08:01]
Second command: /FH debate — models argue multiple rounds over a specific claim (pulled from the DuckDB announcement, about running DuckDB as a server via quack/connect). This is a custom multi-round agentic workflow built into his PI coding agent, not available out of the box in any commercial tool.
[00:09:01] Round one: Fable says the claim is wrong; Gemini 3.7 Flash says keep treating DuckDB as a primary embedded analytical engine; Deepseek V4 Pro says pilot only, not production default. Fable's single debate response cost 15 cents.
[00:10:00] Round two: agents receive each other's opinions and respond. Positions barely move — models converge toward "credible, worth piloting, but not yet a default." Agent names (rune/flux/drift-style aliases) hide the underlying model identity from peer models to prevent sabotage/competition behavior he's observed emerging naturally.
[00:11:01] Final round: closing statements, all reject the original claim as stated, consistent convergence. Total session cost tallied. He frames this as strategic decision-making leverage — for month-long or year-long commitments, it's worth spending pennies of API cost to get multiple independent model perspectives (or counterpoints) before committing.
[00:12:01] Cost comparison callout: Gemini 3.7 Flash and Deepseek V4 Pro are roughly an order of magnitude cheaper than frontier models. Notes a hidden GPT-5.6 pricing tier-jump past 280K tokens (price doubles input, 1.5x output). References his prior week's video on fixing Opus 5's verbosity/over-eagerness via system prompt engineering.
[00:14:00]
Third command: /FH collaborate — most powerful of the three. Prompt: build a working demo of the top three DuckDB release features as Astral UV single-file scripts. All three models propose a plan; two act as "builders," one (the most powerful/expensive model) is designated "architect" and synthesizes the plans into a single execution plan with task IDs (T1/T2/T3), owners, dependencies, and a risk/collision-safety pass — templated via a system prompt approach from his prior prompt-engineering video.
[00:17:01] Architect assigns a todo-list-style task structure with per-task owner + mode + dependencies; agents execute in parallel. He notes this specific 3-model run is for simplicity — the system scales to 5 models ("fusion 5").
[00:18:02] Cost breakdown for the collaborate run: Fable 65¢, Gemini 3.7 Flash 7¢, Deepseek V4 Pro 5¢. Reiterates Gemini 3.7 Flash as likely the fastest, most cost-effective model available right now — argues Google made a deliberate strategic bet to optimize for speed/cost/intelligence rather than chase frontier-only relevance.
[00:19:00] Notes on Deepseek V4 Pro and Quinn 3.8 27B: both "think a lot" (heavy deep-thinking / non-MoE), which hurts local inference speed; Quinn 3.8 benchmarks well on Artificial Analysis but underwhelms in his actual usage due to context-window limitations.
[00:21:01] Core harness-engineering argument: he's still using Claude Code and still paying for the subscription, but argues you cannot do certain things (like this multi-model debate/collaborate workflow) inside someone else's closed agent harness. For engineers pushing into agentic/autonomous work, "the agent harness is the thing to own" — and eventually, owning your own model is the step after that.
[00:23:00] Despite all the new-model coverage, for hardcore agentic engineering work he still orchestrates primarily with Claude Fable 5 over Opus 5 (calls Opus 5 "too hungry," looking for problems that don't exist, though addressable via system prompt engineering from his prior video). Final cost analysis: roughly an order of magnitude cheaper to use Gemini 3.7 Flash or Deepseek V4 Pro vs. a frontier model like Fable — argues that spending real API dollars (not just subscription usage) for multi-model comparison is one of the "easy ways to spend capital to get an advantage," plus API usage carries stronger IP/zero-data-retention protections than some subscription tiers.
[00:26:01] Thesis wrap: "combine compute, don't select compute" is the catchphrase. Beyond the fusion-harness multi-model pattern, he reiterates the channel's ongoing "software factory" thesis — combining agents + code to operate outloop (agents working autonomously on your behalf) rather than inloop (babysitting one agent in a terminal) as the next leverage tier past harness engineering.
[00:28:02] Closing: intelligence explosions will keep happening; positions the channel as a trusted tracking source for model releases and harness patterns; sign-off.