06-reference/research

llm forecasting brier gap tool access

2026-08-02·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
forecasting-calibrationbrier-scoreretrieval-augmentationtargeting-systeminstrumentable-frontier

Tool access buys crowd-level calibration, not expert-level — and it buys it by reading the crowd

The question

"Does the LLM-vs-superforecaster Brier-score gap narrow in domains where the model has tool-access to live ground truth (search, fleet data)? Is the instrumentable frontier expanding fast enough to re-price the residual conviction gap?"

Follow-up #1 from [[2026-06-12-agentic-targeting-conviction-calibrated-confidence]], auto-promoted 2026-06-13. The load-bearing angle is the founder's targeting-system filter: if instrumentation is now cheap enough to make agent forecasts trustworthy, the set of bets RDCO can take changes. If not, the residual stays human and the positioning stays as drafted.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

The answer to the literal question is yes, and by a lot — the retrieval ablation is the cleanest number in this whole research cluster. Giving a model live search moved it from .206 to .179 on 914 contamination-controlled questions, while a year of frontier model progress buys roughly .015 on ForecastBench. Even allowing that those are different rulers and cannot be directly subtracted, the order of magnitude is worth internalizing: one good retrieval install is worth something like a generation or two of model upgrades. For RDCO's own machinery that is the operational lever. The highest-value change to the /verify-* stack and the conviction pipeline is not "wait for a better model," it is "wire the critic to live ground truth for the specific claim it is checking." That is a build we control and it dominates waiting.

The answer to the second question — is the frontier expanding fast enough to re-price the residual — is no, and for a structural reason rather than a timing one. The mechanism producing the gain is retrieval over public news plus, at the frontier, anchoring on an existing human crowd forecast. Both inputs exist only in domains that are already instrumented. A liquid prediction market, a dense news corpus, and a resolvable question with a date on it are the preconditions, not the outputs. So the instrumentable frontier is expanding fastest exactly where instrumentation already lived, and it is not advancing into the un-instrumented residual at all. The residual keeps its shape. The parent brief's boundary holds; what changed is that the already-instrumented band got substantially cheaper to work in, and RDCO should harvest that band aggressively rather than reprice the fuzz.

This sharpens the MAC and agent-deployer pitch rather than softening it. The evidence-backed version is now more concrete than "we narrow the fuzz": the measured payoff of instrumentation is a step-change in calibration, and the measured payoff of model progress alone is incremental. MAC's product is the retrieval-and-resolution layer — turning a client's fuzzy judgment call into a question with a resolution date, an evidence corpus the agent can search, and an observable outcome. That is the thing that moves Brier. Selling "our agent is smarter" is selling the .015; selling "we install the feedback loop" is selling the .027. The credibility moat from the parent brief also gets a new plank: any vendor citing "AI now matches superforecasters" is very likely quoting a difficulty-adjusted score computed across non-overlapping question sets with no published standard error, on a benchmark whose own authors put a 22-month confidence interval on the parity date.

For the investing track, this brief should reduce rather than increase confidence in agent-issued forward calls. The chip-fab and memory capital-cycle thesis is close to a worst case for the mechanism that produces gap-narrowing: long horizon (position-scale, not weeks), low N of decision points, no liquid resolution market to anchor to, and a thesis that resolves through a slow sequence of ambiguous confirmations rather than a dated binary. The gains documented here concentrate in short-horizon, news-covered, dated binaries. Nothing here licenses letting a model size a position. The prior conclusions stand unchanged: the [[brier-score]] 0.12 held-out gate remains the deploy criterion, and per [[2026-06-16-multi-agent-ensembles-conviction-calibration]] a panel's disagreement spread is the usable signal while its point estimate is not. What is newly justified is instrumenting the thesis better — pushing more of the capital-cycle call into dated, resolvable sub-questions the agent can retrieve against, which is the same "expand the instrumentable frontier" move, applied inward.

Why this is in the vault

It answers the first open follow-up from [[2026-06-12-agentic-targeting-conviction-calibrated-confidence]] and settles whether the founder's targeting-system prioritization filter needs revision: it does not, because tool access expands capability inside already-instrumented domains rather than annexing the un-instrumented residual. It also retires the unverified "green tree matched superforecasters" claim carried in [[2026-05-23-innermost-loop-singularity-forecasts-itself]], which that note explicitly asked a later brief to check, and it supplies the calibration caveat (difficulty-adjusted, non-overlapping question sets, no published SE) that any RDCO artifact citing ForecastBench must carry.

Open follow-ups

Related

Sources