Tool access buys crowd-level calibration, not expert-level — and it buys it by reading the crowd
The question
"Does the LLM-vs-superforecaster Brier-score gap narrow in domains where the model has tool-access to live ground truth (search, fleet data)? Is the instrumentable frontier expanding fast enough to re-price the residual conviction gap?"
Follow-up #1 from [[2026-06-12-agentic-targeting-conviction-calibrated-confidence]], auto-promoted 2026-06-13. The load-bearing angle is the founder's targeting-system filter: if instrumentation is now cheap enough to make agent forecasts trustworthy, the set of bets RDCO can take changes. If not, the residual stays human and the positioning stays as drafted.
What we already know (from the vault)
- The parent brief [[2026-06-12-agentic-targeting-conviction-calibrated-confidence]] found single frontier LLMs trailing expert superforecasters roughly 6x on Brier (o3 0.135 vs expert median 0.023), systematically overconfident where stakes are highest, and moving human commitment only via "confidence alignment" contamination rather than earned conviction. It concluded conviction in un-instrumented domains stays human, and named real-time observable feedback as the one lever that fixes miscalibrated reliance.
- [[2026-06-16-multi-agent-ensembles-conviction-calibration]] tested the other escape hatch and found a ceiling: ensembles and debate reach crowd calibration (Brier ~0.20 on that benchmark), not superforecaster, and do not cure overconfidence. Aggregation rescues weak pools; it does not beat a strong single judge.
- [[brier-score]] supplies the discipline this brief has to hold itself to: "the 'good' Brier number depends on how hard the prediction problem is," always report N, and always name the baseline. It also records RDCO's own operational gate — no strategy deploys real capital until it beats Brier 0.12 on held-out data.
- [[2026-05-23-innermost-loop-singularity-forecasts-itself]] carried a secondhand newsletter claim that a Google DeepMind model ("green tree") took the ForecastBench top spot and "matched superforecaster performance for the first time." That vault note itself flagged the claim as needing verification against the actual leaderboard before being treated as actionable. This brief is that follow-up, and the claim is still not verified (see below).
- [[2026-04-24-targeting-system]] frames the whole question: the agentic targeting system works where outcomes are well-defined, and the residual fuzz is where conviction lives.
What the web says
- The cleanest causal evidence that tool access narrows the gap is a retrieval ablation, and it is large. Halawi et al., Approaching Human-Level Forecasting with Language Models (NeurIPS 2024, arXiv 2402.18563): the full retrieval-augmented system scored Brier .179 ± .003 on a test set of 914 binary questions. Strip retrieval and fine-tuning and the same base model scores .206 — essentially the pre-trained zero-shot baseline (.208). Retrieval plus fine-tuning is worth roughly .027 Brier, which is most of the distance from "knows only its priors" to "near the human crowd."
- But it closes the gap to the crowd, not to superforecasters, and it does not close it fully. The human crowd aggregate on the same set scored .149 (reported as .146 ± .002 in the paper's Table 5), so the system still trails by ~.03. The system beats the crowd only in a jointly selective setting — crowd probability in [.3,.7], early retrieval dates, and ≥5 relevant articles retrieved — covering 22% of forecasts / 43% of questions, and there by ">1.5 standard errors." On the ≥5-articles condition alone the crowd still wins (.175 system vs .143 crowd). The authors explicitly caution against reading subcategory results as signal given smaller N, and note the system underperforms the crowd precisely where the crowd is confident.
- Contamination control in this specific paper is genuinely good. All test questions resolve June 1, 2023 or later, after the base models' training cutoff; train and validation questions all resolved before that date; questions spanning the boundary were discarded. The news corpus was restricted to articles predating the retrieval date. This is the well-designed end of the literature and the results should be weighted accordingly.
- On the live benchmark, the frontier is close but the headline number is a modelled estimate, not a head-to-head. The Forecasting Research Institute reports that as of January 29, 2026 superforecasters held #1 on ForecastBench, ahead of the best LLM entries by 0.017 Brier points, with xAI's Grok 4.20 (Preview) and Cassi's
ensemble_2_crowdadjtied at #2 (FRI). Both top entries are tool-augmented: Grok ran X search, web search and a Python REPL, averaging eight forecasts; Cassi used Tavily retrieval plus a model ensemble. No question count, standard error, or confidence interval on that 0.017 gap is reported. - The gap is computed across non-overlapping question sets. FRI states the initial superforecaster sample and the latest models "have zero overlapping questions," and the ranking uses a difficulty-adjusted Brier: a two-way fixed-effects decomposition into forecaster-ability and question-difficulty terms (ForecastBench updated methodology). The method's stated assumption is that "forecaster quality remains stable over time." Its validation simulation used 141 forecasters over 473 resolved binary questions (422 machine-generated dataset questions from ACLED/FRED/DBnomics/Wikipedia/Yahoo Finance, 51 market questions), with perfect overlap — but that is the simulation, not the live leaderboard, and the count of superforecasters inside the 141 is not stated. No SE or significance test on any model-vs-superforecaster gap appears in that document.
- ForecastBench does guard contamination structurally, and says so. Questions resolve in the future relative to submission, and FRI explicitly excludes models more than a year past their training cutoff from the difficulty-estimation regression on the grounds that "such contamination would bias our difficulty estimates." There is also a 50-day waiting period before new models enter the leaderboard.
- Projected parity has a two-year confidence interval. FRI extrapolates LLM-superforecaster parity in November 2026, 95% CI January 2026 – November 2027, from a measured improvement of ~0.015 Brier per year (Claude 3.5 Sonnet 0.117 in Oct 2024 → Grok 4.20 Preview 0.102 in Oct 2025). The trend is real; the date is not knowable from this data.
- Unverified in this run: search snippets surfaced a later FRI post titled "AI models have likely reached parity with superforecasters on ForecastBench" and a July 16, 2026 leaderboard state in which several systems were described as statistically indistinguishable from superforecaster accuracy. I did not fetch either source, so I am not asserting those numbers. The Good Judgment rebuttal, "What ForecastBench Doesn't Measure (Yet)," returned HTTP 403 and was not retried. The vault's "green tree topped ForecastBench" claim remains unverified against a primary source.
Convergences and contradictions
- Convergence, with a correction to the parent brief's headline number. Tool access does narrow the gap, and materially. But the ceiling the ensembles brief found is the same ceiling here: retrieval buys crowd calibration. Two independent literatures — ensembling and retrieval — both land at the human crowd and stop short of expert superforecasters. The parent brief's "~6x gap" is stale as a figure while intact as a claim.
- The three "gaps" in this cluster are measured with incommensurate rulers. 0.135 vs 0.023 (parent brief's source), .179 vs .149 (Halawi, N=914), and 0.017 difficulty-adjusted (ForecastBench). The superforecaster baseline itself ranges from 0.023 to ~0.10 across these. Per [[brier-score]], that spread is question difficulty, not forecaster skill. Any narrative of the form "the gap shrank from X to Y" that crosses these benchmarks is measuring the question sets, not the models. That includes the tempting framing of this very brief, and it is the single most likely way this literature gets oversold.
- A contradiction that matters more than the headline. The #2 ForecastBench entry is Cassi's
ensemble_2_crowdadj, and FRI notes the crowd-adjusted variant beats Cassi's own un-adjusted ensemble by nearly 0.01 Brier. Roughly half the remaining gap-closure at the frontier appears to come from anchoring on the human crowd, not from independent model reasoning. The systems that most nearly match superforecasters are partly borrowing human forecasts as an input. That is a real capability, but it is not the capability the question is asking about. - The "fleet data" half of the question is unanswered by the literature. Every result found here is grounded in public news retrieval or prediction-market prices. I found no published forecasting evaluation grounded in private operational telemetry. Treat "search" and "fleet data" as separate cases; only the first has evidence.
Synthesis for RDCO
The answer to the literal question is yes, and by a lot — the retrieval ablation is the cleanest number in this whole research cluster. Giving a model live search moved it from .206 to .179 on 914 contamination-controlled questions, while a year of frontier model progress buys roughly .015 on ForecastBench. Even allowing that those are different rulers and cannot be directly subtracted, the order of magnitude is worth internalizing: one good retrieval install is worth something like a generation or two of model upgrades. For RDCO's own machinery that is the operational lever. The highest-value change to the /verify-* stack and the conviction pipeline is not "wait for a better model," it is "wire the critic to live ground truth for the specific claim it is checking." That is a build we control and it dominates waiting.
The answer to the second question — is the frontier expanding fast enough to re-price the residual — is no, and for a structural reason rather than a timing one. The mechanism producing the gain is retrieval over public news plus, at the frontier, anchoring on an existing human crowd forecast. Both inputs exist only in domains that are already instrumented. A liquid prediction market, a dense news corpus, and a resolvable question with a date on it are the preconditions, not the outputs. So the instrumentable frontier is expanding fastest exactly where instrumentation already lived, and it is not advancing into the un-instrumented residual at all. The residual keeps its shape. The parent brief's boundary holds; what changed is that the already-instrumented band got substantially cheaper to work in, and RDCO should harvest that band aggressively rather than reprice the fuzz.
This sharpens the MAC and agent-deployer pitch rather than softening it. The evidence-backed version is now more concrete than "we narrow the fuzz": the measured payoff of instrumentation is a step-change in calibration, and the measured payoff of model progress alone is incremental. MAC's product is the retrieval-and-resolution layer — turning a client's fuzzy judgment call into a question with a resolution date, an evidence corpus the agent can search, and an observable outcome. That is the thing that moves Brier. Selling "our agent is smarter" is selling the .015; selling "we install the feedback loop" is selling the .027. The credibility moat from the parent brief also gets a new plank: any vendor citing "AI now matches superforecasters" is very likely quoting a difficulty-adjusted score computed across non-overlapping question sets with no published standard error, on a benchmark whose own authors put a 22-month confidence interval on the parity date.
For the investing track, this brief should reduce rather than increase confidence in agent-issued forward calls. The chip-fab and memory capital-cycle thesis is close to a worst case for the mechanism that produces gap-narrowing: long horizon (position-scale, not weeks), low N of decision points, no liquid resolution market to anchor to, and a thesis that resolves through a slow sequence of ambiguous confirmations rather than a dated binary. The gains documented here concentrate in short-horizon, news-covered, dated binaries. Nothing here licenses letting a model size a position. The prior conclusions stand unchanged: the [[brier-score]] 0.12 held-out gate remains the deploy criterion, and per [[2026-06-16-multi-agent-ensembles-conviction-calibration]] a panel's disagreement spread is the usable signal while its point estimate is not. What is newly justified is instrumenting the thesis better — pushing more of the capital-cycle call into dated, resolvable sub-questions the agent can retrieve against, which is the same "expand the instrumentable frontier" move, applied inward.
Why this is in the vault
It answers the first open follow-up from [[2026-06-12-agentic-targeting-conviction-calibrated-confidence]] and settles whether the founder's targeting-system prioritization filter needs revision: it does not, because tool access expands capability inside already-instrumented domains rather than annexing the un-instrumented residual. It also retires the unverified "green tree matched superforecasters" claim carried in [[2026-05-23-innermost-loop-singularity-forecasts-itself]], which that note explicitly asked a later brief to check, and it supplies the calibration caveat (difficulty-adjusted, non-overlapping question sets, no published SE) that any RDCO artifact citing ForecastBench must carry.
Open follow-ups
- Does the ForecastBench superforecaster-vs-model gap survive a head-to-head round on identical questions, and has anyone published a standard error on it? The 0.017 figure is a fixed-effects estimate across disjoint question sets; without an SE it is not possible to say whether "statistically indistinguishable" reflects model skill or the benchmark's noise floor.
- How much of frontier LLM forecasting skill is borrowed from the human crowd? Cassi's crowd-adjusted ensemble beats its own un-adjusted version by ~0.01. Is there a published leaderboard variant with crowd-anchoring ablated, and what does the ranking look like without it?
- Is there any published forecasting evaluation grounded in private operational telemetry (fleet, sensor, internal BI) rather than public news retrieval? This is the literal "fleet data" half of the question and the search found nothing.
- Does the retrieval gain hold at long horizons? The documented gains cluster on short-horizon, news-covered questions where retrieval is close to the answer. Is there an evaluation stratified by resolution horizon (weeks vs 6-24 months) that would tell us whether any of this transfers to position-horizon investing calls?
- Where is the news-coverage threshold? Halawi's system needed ≥5 relevant retrieved articles to reach its better scores. What happens in the low-coverage tail, and is there a measurable coverage floor below which retrieval-augmented forecasts revert to base-model calibration? This is the practical spec for whether a given client domain is instrumentable.
- What are Good Judgment's specific objections in "What ForecastBench Doesn't Measure (Yet)" (403 in this run), and do they survive scrutiny? The incumbent superforecaster vendor is a motivated critic, which makes the argument worth reading and worth discounting.
Related
- [[2026-06-12-agentic-targeting-conviction-calibrated-confidence]]
- [[2026-06-16-multi-agent-ensembles-conviction-calibration]]
- [[brier-score]]
- [[2026-04-24-targeting-system]]
- [[2026-05-23-innermost-loop-singularity-forecasts-itself]]
- [[2026-06-18-probability-aggregation-scoring-rules-panel]]
- [[2026-06-20-conviction-score-binary-collapse-point]]
- [[2026-04-19-commoncog-the-forecasting-series]]
- [[2026-05-27-markov-equities-pipeline-spec]]
feedback_targeting_system_prioritization_filter(memory)
Sources
- Vault: [[2026-06-12-agentic-targeting-conviction-calibrated-confidence]] —
~/rdco-vault/06-reference/research/2026-06-12-agentic-targeting-conviction-calibrated-confidence.md(parent brief; the ~6x gap, overconfidence, confidence-alignment contamination) - Vault: [[2026-06-16-multi-agent-ensembles-conviction-calibration]] —
~/rdco-vault/06-reference/research/2026-06-16-multi-agent-ensembles-conviction-calibration.md(ensembles reach crowd, not expert, calibration) - Vault: [[brier-score]] —
~/rdco-vault/06-reference/concepts/brier-score.md(baseline discipline, report-the-N rule, RDCO's 0.12 deploy gate) - Vault: [[2026-04-24-targeting-system]] —
~/rdco-vault/06-reference/concepts/2026-04-24-targeting-system.md(implicit vs agentic split; the residual) - Vault: [[2026-05-23-innermost-loop-singularity-forecasts-itself]] —
~/rdco-vault/06-reference/2026-05-23-innermost-loop-singularity-forecasts-itself.md(the unverified "green tree" ForecastBench claim this brief was asked to check) - Vault: [[2026-05-27-markov-equities-pipeline-spec]] —
~/rdco-vault/01-projects/investing/2026-05-27-markov-equities-pipeline-spec.md(the investing-conviction surface this bears on) - Web: Halawi et al., Approaching Human-Level Forecasting with Language Models — https://arxiv.org/pdf/2402.18563 (retrieval ablation .206 → .179 ± .003, N=914, crowd .149; contamination-controlled post-June-2023 test set). PDF did not parse via WebFetch; text extracted locally from the saved file.
- Web: Forecasting Research Institute, LLMs Are Closing the Gap on Human Superforecasters — https://forecastingresearch.substack.com/p/llms-are-closing-the-gap-on-human (0.017 difficulty-adjusted gap as of 2026-01-29; parity extrapolation Nov 2026, 95% CI Jan 2026 – Nov 2027; tool configurations of top entries)
- Web: ForecastBench: Updated Ranking Methodology (Kucinskas) — https://www.forecastbench.org/assets/pdfs/forecastbench_updated_methodology.pdf (two-way fixed-effects difficulty adjustment, non-overlapping question sets, contamination exclusion rule, 141 forecasters / 473 questions in the validation simulation). PDF did not parse via WebFetch; text extracted locally from the saved file.
- Web (SKIPPED — HTTP 403, not retried): Good Judgment, What ForecastBench Doesn't Measure (Yet) — https://goodjudgment.com/what-forecastbench-doesnt-measure/
- Web (NOT FETCHED — surfaced in search snippets only, numbers not asserted in this brief): FRI, AI models have likely reached parity with superforecasters on ForecastBench — https://forecastingresearch.substack.com/p/ai-models-have-likely-reached-parity