06-reference/research

llm as judge surface heuristic bias

2026-09-25·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
llm-as-judgecritic-layersurface-heuristicsstyle-biasverification

Style bias dominates, agreement is mostly rubric echo, and more voters is the wrong lever

The question

Verbatim: "What does the 2025-2026 evidence base say about when LLM-as-judge fails due to surface heuristics (formatting, fluency, tone over correctness), and which prompting or multi-voter techniques most reliably detect or correct for this bias?"

Context: RDCO's [[2026-07-03-verify-skills-pass-disagreement-audit]] named surface-over-substance as the #1 structural exposure of the verify-* critic family and cited arXiv 2603.11027 for it. This brief checks that citation against the primary source and asks what the last two quarters actually established about fixes.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

The single change with the best evidence behind it is not a new critic — it is splitting every existing critic into a verifier pass and a judge pass, and forbidding the judge to score anything the verifier can decide. The 62% result says a rubric is doing most of the work that RDCO currently believes judgment is doing, and the JudgeBench-class results say the judgment left over is near-chance precisely on correctness. That combination means every axis RDCO can make mechanical should stop being an LLM axis today. Concretely, the next audit-*.py-shaped script should resolve every citation in a vault note (does the URL or arXiv ID resolve, does the returned title match the cited title, does the cited author match), assert every [[wikilink]] target exists on disk, and assert every numeral in a synthesis paragraph appears in a cited source. [[2026-07-19-constitution-or-collapse-citation-verification]] is the existence proof that this class of defect is real in RDCO's corpus and that a presence-checking judge waves it through. This is also the cheapest work in the brief: it is a script, it is deterministic, and it never gets talked into a PASS by a well-formatted paragraph.

The second change is style-blinding on the substance axes, and it is a two-pass split rather than a prompt tweak. Style bias measured at 0.10-0.76 against position bias at or below 0.04 is a clear instruction about where to spend. Telling a critic "do not be swayed by formatting" is the weakest form of the fix, because the surface signal is in the artifact itself. The stronger form: for verify-vault-write, verify-strategic-output and verify-dispatch, render a de-styled copy of the artifact (headings flattened, bold and bullets stripped, confident framing left intact but structure removed) and hand that to the substance critic, while the format axes get the styled original in a separate pass with no substance authority. The critics where format is genuinely the object under test (verify-pdf-output, design-critic, station-critic on visual axes) are exempt by construction, which is fine: those are honest surface checks that never claimed to score substance. The dishonest case is a single pass that reads a polished artifact and returns one verdict covering both.

The third change concerns the iterate loop, and it is a rule, not a build. RDCO's fresh-eyes doctrine ([[2026-05-19-verification-as-independent-worker-pattern]]) already says the critic must be zero-context relative to the producer. The gap is the re-review: when an artifact comes back after a FAIL, the natural implementation hands the critic the prior verdict and the repair diff so it can check the fix. arXiv 2608.16003's title alone is enough to make that suspicious, and the mechanism is obvious without reading it: a judge that has just watched a repair is primed to accept it. The rule to write into the verify-* SKILL.md family: on re-review, spawn a new critic that has never seen the prior verdict, the defect list, or the repair history — only the current artifact and the rubric. Re-review cost is one more subagent, which is nothing against a false PASS on a strategic output.

On multi-voter, the honest answer is: RDCO already has the right shape and should stop shopping for a better one. The per-axis symmetric fan-out in station-critic matches what the 2026 evidence supports, and both the vendor-diversity arm ([[2026-07-30-multi-llm-critic-ensembles-differentiated-rubrics]]) and the role-asymmetric debate arm (2608.30373) are now negative results. The remaining work is population, not architecture: three axis fragments exist and the entire verify-* family still runs one monolithic rubric per critic. And none of it is measurable until the verdict log exists. That log was named the binding constraint on 2026-07-03, I verified it is still absent on 2026-09-25, and every proposal in this brief reduces to an unfalsifiable opinion without it. The detection experiment that finally makes the bias visible is small and specific: take ten artifacts RDCO already PASSed, produce a twin of each with one injected substantive error and better polish than the original (tighter headings, more confident prose, denser formatting), run the current critics blind on both, and count PASSes on the polished-wrong twins. A high PASS rate converts "we suspect surface bias" into a number. That is the July audit's unrun H1 with a concrete design attached.

Why this is in the vault

This closes the citation-verification question left hanging by [[2026-07-03-verify-skills-pass-disagreement-audit]] (arXiv 2603.11027 is real and says what the audit said it says), and it redirects the audit's own H3 — pilot a cross-model critic — which the 2026 evidence now contraindicates in favor of a verifier/judge split and style-blinding on the same model. It is the input to the next concrete edits to the verify-vault-write, verify-strategic-output and verify-dispatch SKILL.md files and to the still-unbuilt verdict log.

Open follow-ups

Related

Sources

Vault:

Web (fetched and verified this run):

Web (search-summary level only, NOT fetched — treat as unverified):

Paywalls / access notes: none hit this run. Full PDFs of 2604.23178 and 2608.30373 were not retrieved; findings above come from their abstract pages only.