Style bias dominates, agreement is mostly rubric echo, and more voters is the wrong lever
The question
Verbatim: "What does the 2025-2026 evidence base say about when LLM-as-judge fails due to surface heuristics (formatting, fluency, tone over correctness), and which prompting or multi-voter techniques most reliably detect or correct for this bias?"
Context: RDCO's [[2026-07-03-verify-skills-pass-disagreement-audit]] named surface-over-substance as the #1 structural exposure of the verify-* critic family and cited arXiv 2603.11027 for it. This brief checks that citation against the primary source and asks what the last two quarters actually established about fixes.
What we already know (from the vault)
- The exposure was already diagnosed, and the diagnosis was explicitly unfalsifiable. [[2026-07-03-verify-skills-pass-disagreement-audit]] found zero logged founder-vs-PASS disagreements in a 60-day window and called that an instrumentation gap, not a clean bill of health: "zero overrides" and "critics are well-calibrated" are indistinguishable without a verdict log. It ranked surface-over-substance first among blind spots because verify-pdf-output scores 12 layout invariants and nothing about whether the memo's argument holds, and verify-vault-write scores whether an RDCO mapping is present, not whether it is right.
- Adding vendors does not buy independence. [[2026-07-30-multi-llm-critic-ensembles-differentiated-rubrics]] found RDCO's own seeded-defect benchmark produced zero unique catches from Codex or Grok, and that the outside literature agrees at much larger N (roughly 2.0-2.4 effective votes out of nine judges). Its conclusion: treat rubric-sharding as a coverage upgrade, not an independence fix, and keep external models unwired.
- The two-tier architecture is already correct in principle. [[2026-05-20-verify-stack-two-gate-pass-fail-architecture]] pairs a deterministic tier (the 13-invariant
audit-newsletter-outputs.py, zero LLM calls) with the LLM-judge tier. The deterministic tier is the part that cannot be talked into a PASS by polish. - RDCO has already caught a real surface-heuristic miss in production. [[2026-07-19-constitution-or-collapse-citation-verification]] found a vault note citing a real paper under a flatly wrong author attribution, with a described failure mode that appears nowhere in the source. A critic reading for citation presence passes that. A verifier resolving the citation does not.
- Verified on disk today (2026-09-25), not inferred:
~/.claude/state/improve-cron-log.mdandflash-review-log.jsonlstill do not exist twelve weeks after the audit named them the binding constraint; onlybehavior-critic-log.jsonlwas ever stood up.~/rdco-vault/01-projects/skill-pipelines/axes/still holds exactly three fragments (image-coherence, novelty, syllable-count), none ported to verify-. No verify- or station-criticSKILL.mdcontains any swap, position-bias, or blinding language.
What the web says
- The internal citation checks out. arXiv 2603.11027 resolves to "Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge" (Song, Zheng, Xu; submitted 2026-03-11) — verified by fetching the abstract page this run. Across 105,600 evaluation instances it reports that model-level agreement masks fragile sample-level agreement, and that shared rubric structure accounts for 62% of total agreement. Its remedy is MERG, knowledge-driven rubric generation that injects domain expertise instead of generic criteria. (https://arxiv.org/abs/2603.11027)
- Style bias is roughly an order of magnitude larger than position bias. "Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines" (Soumik; submitted 2026-04-25) evaluates nine debiasing strategies and reports style bias in the 0.10-0.76 range while position bias stays at or below 0.04. Debiasing lifted agreement by +4.5 to +11.5 percentage points depending on model, with a best configuration at 71.0% agreement (kappa 0.549). Honesty flag: the abstract page did not enumerate which of the nine strategies each delta belongs to, so I cannot attribute any specific number to any specific technique. (https://arxiv.org/abs/2604.23178)
- Multi-agent debate made judging worse, and the culprit was role asymmetry. "Beyond Consensus: Downward Bias and Role Asymmetry in Multi-Agent LLM Judges" (Song, Kim, Eo, Park; submitted 2026-08-31) found the single-judge baseline achieved the strongest human alignment on average across six models, with multi-agent debate degrading it. Assigning one agent a "strict judge" role produced systematic downward bias that consensus failed to correct, landing well past the arithmetic midpoint ("strict-stance dominance"); symmetric-role debate largely recovered baseline. (https://arxiv.org/abs/2608.30373)
- Judges are near-chance when correctness is the ground truth. JudgeBench reports GPT-4o at roughly 56.6% on objectively-checkable reasoning, math, and code pairs against a 50% random baseline, with several fine-tuned judge models scoring below random. Search-summary level only, not fetched this run. (https://arxiv.org/pdf/2410.12784)
- The strongest reported position-bias control is double-ordering, not voting: run A-then-B and B-then-A and require both verdicts to agree. Practitioner sources also converge on decomposing coarse rubrics into discrete atomic checks and validating each against a human baseline until correlation clears ~0.85, and note that verbosity bias varies by model family, so a calibration tuned on one provider does not transfer. Search-summary level, practitioner sources, not primary research.
- Two 2026 titles worth pulling next round, surfaced but not read: "Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency" (arXiv 2608.16003) and "A Judge Should Know What Changed: Construct Validity for LLM-as-a-Judge Evaluation" (arXiv 2608.24419). Both bear directly on RDCO's iterate loop. Titles only; no claim made about their contents.
Convergences and contradictions
- Convergence, and it is the sharpest thing in this brief. The 62% figure and RDCO's own zero-unique-catches benchmark are the same finding from two directions: when judges share a rubric, their agreement is mostly rubric echo rather than independent confirmation. RDCO's verify-* family shares one rubric and one model and one producer lineage, which is the maximally-correlated configuration. Agreement between two RDCO critics on the same artifact carries close to no information.
- Contradiction with the vault's own default instinct. The July audit's H3 proposed piloting a cross-model or differently-prompted critic, and the general reflex when a judge looks unreliable is to add voters. Both 2608.30373 and the effective-votes literature say that is the weakest lever available, and that a role-asymmetric panel (one agent told to be strict) is actively worse than a single judge. RDCO's station-critic per-axis fan-out is symmetric, which is the supported shape; any future "harsh critic vs advocate" design is contraindicated by the 2026 evidence.
- Transfer caveat the literature does not flag and RDCO must. Most 2025-2026 debiasing work measures pairwise preference judging. RDCO's critics are pointwise gates: one artifact, PASS or FAIL. Position swapping and jury voting are pairwise instruments and largely do not apply. Style bias, self-preference, and fluency inflation transfer fully. So the headline import is narrow: spend nothing on position-bias controls, spend everything on style-blinding and on moving checks out of the judge.
Synthesis for RDCO
The single change with the best evidence behind it is not a new critic — it is splitting every existing critic into a verifier pass and a judge pass, and forbidding the judge to score anything the verifier can decide. The 62% result says a rubric is doing most of the work that RDCO currently believes judgment is doing, and the JudgeBench-class results say the judgment left over is near-chance precisely on correctness. That combination means every axis RDCO can make mechanical should stop being an LLM axis today. Concretely, the next audit-*.py-shaped script should resolve every citation in a vault note (does the URL or arXiv ID resolve, does the returned title match the cited title, does the cited author match), assert every [[wikilink]] target exists on disk, and assert every numeral in a synthesis paragraph appears in a cited source. [[2026-07-19-constitution-or-collapse-citation-verification]] is the existence proof that this class of defect is real in RDCO's corpus and that a presence-checking judge waves it through. This is also the cheapest work in the brief: it is a script, it is deterministic, and it never gets talked into a PASS by a well-formatted paragraph.
The second change is style-blinding on the substance axes, and it is a two-pass split rather than a prompt tweak. Style bias measured at 0.10-0.76 against position bias at or below 0.04 is a clear instruction about where to spend. Telling a critic "do not be swayed by formatting" is the weakest form of the fix, because the surface signal is in the artifact itself. The stronger form: for verify-vault-write, verify-strategic-output and verify-dispatch, render a de-styled copy of the artifact (headings flattened, bold and bullets stripped, confident framing left intact but structure removed) and hand that to the substance critic, while the format axes get the styled original in a separate pass with no substance authority. The critics where format is genuinely the object under test (verify-pdf-output, design-critic, station-critic on visual axes) are exempt by construction, which is fine: those are honest surface checks that never claimed to score substance. The dishonest case is a single pass that reads a polished artifact and returns one verdict covering both.
The third change concerns the iterate loop, and it is a rule, not a build. RDCO's fresh-eyes doctrine ([[2026-05-19-verification-as-independent-worker-pattern]]) already says the critic must be zero-context relative to the producer. The gap is the re-review: when an artifact comes back after a FAIL, the natural implementation hands the critic the prior verdict and the repair diff so it can check the fix. arXiv 2608.16003's title alone is enough to make that suspicious, and the mechanism is obvious without reading it: a judge that has just watched a repair is primed to accept it. The rule to write into the verify-* SKILL.md family: on re-review, spawn a new critic that has never seen the prior verdict, the defect list, or the repair history — only the current artifact and the rubric. Re-review cost is one more subagent, which is nothing against a false PASS on a strategic output.
On multi-voter, the honest answer is: RDCO already has the right shape and should stop shopping for a better one. The per-axis symmetric fan-out in station-critic matches what the 2026 evidence supports, and both the vendor-diversity arm ([[2026-07-30-multi-llm-critic-ensembles-differentiated-rubrics]]) and the role-asymmetric debate arm (2608.30373) are now negative results. The remaining work is population, not architecture: three axis fragments exist and the entire verify-* family still runs one monolithic rubric per critic. And none of it is measurable until the verdict log exists. That log was named the binding constraint on 2026-07-03, I verified it is still absent on 2026-09-25, and every proposal in this brief reduces to an unfalsifiable opinion without it. The detection experiment that finally makes the bias visible is small and specific: take ten artifacts RDCO already PASSed, produce a twin of each with one injected substantive error and better polish than the original (tighter headings, more confident prose, denser formatting), run the current critics blind on both, and count PASSes on the polished-wrong twins. A high PASS rate converts "we suspect surface bias" into a number. That is the July audit's unrun H1 with a concrete design attached.
Why this is in the vault
This closes the citation-verification question left hanging by [[2026-07-03-verify-skills-pass-disagreement-audit]] (arXiv 2603.11027 is real and says what the audit said it says), and it redirects the audit's own H3 — pilot a cross-model critic — which the 2026 evidence now contraindicates in favor of a verifier/judge split and style-blinding on the same model. It is the input to the next concrete edits to the verify-vault-write, verify-strategic-output and verify-dispatch SKILL.md files and to the still-unbuilt verdict log.
Open follow-ups
- Read arXiv 2608.16003 ("Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency") against RDCO's iterate loop: does the leniency shift survive a fresh instantiation, or only a same-session one?
- Read arXiv 2608.24419 ("A Judge Should Know What Changed") — if construct validity requires the judge to see the diff, it directly contradicts the zero-context re-review rule proposed above, and the conflict needs resolving before either is written into a SKILL.md.
- Enumerate which of the nine debiasing strategies in arXiv 2604.23178 produced which deltas; the abstract page withheld the mapping and the full PDF was not fetched this run.
- Does style-blinding measurably change verdicts on RDCO's own corpus, or is the de-styled artifact simply harder to judge at all (higher FAIL rate without higher catch rate)?
- Is the 0.10-0.76 style-bias range a pairwise-only artifact, or does it reproduce in pointwise PASS/FAIL gating? No source found this run measures pointwise style bias directly.
- Stand up the verdict log (
improve-cron-log.md/ per-axisflash-review-log.jsonl) — carried unresolved from 2026-07-03 and re-verified absent 2026-09-25. - Build the deterministic citation-resolution verifier as the first non-LLM axis for verify-vault-write; scope and owner unset.
Related
- [[2026-07-03-verify-skills-pass-disagreement-audit]]
- [[2026-07-30-multi-llm-critic-ensembles-differentiated-rubrics]]
- [[2026-05-20-verify-stack-two-gate-pass-fail-architecture]]
- [[2026-05-19-verification-as-independent-worker-pattern]]
- [[2026-05-22-reward-hacking-patterns-llm-critic-systems]]
- [[2026-05-19-cai-critic-graduation-per-axis-threshold]]
- [[2026-07-19-constitution-or-collapse-citation-verification]]
- [[2026-06-14-open-ended-output-verifier]]
- [[2026-05-23-improve-cron-design-spec]]
Sources
Vault:
- ~/rdco-vault/06-reference/research/2026-07-03-verify-skills-pass-disagreement-audit.md
- ~/rdco-vault/06-reference/research/2026-07-30-multi-llm-critic-ensembles-differentiated-rubrics.md
- ~/rdco-vault/02-sops/2026-05-20-verify-stack-two-gate-pass-fail-architecture.md
- ~/rdco-vault/02-sops/2026-05-19-verification-as-independent-worker-pattern.md
- ~/rdco-vault/06-reference/research/2026-05-22-reward-hacking-patterns-llm-critic-systems.md
- ~/rdco-vault/06-reference/research/2026-05-19-cai-critic-graduation-per-axis-threshold.md
- ~/rdco-vault/06-reference/research/2026-07-19-constitution-or-collapse-citation-verification.md
- ~/rdco-vault/06-reference/research/2026-06-14-open-ended-output-verifier.md
- ~/rdco-vault/06-reference/research/2026-05-23-improve-cron-design-spec.md
- Filesystem state checked directly 2026-09-25: ~/.claude/state/, ~/rdco-vault/01-projects/skill-pipelines/axes/, ~/.claude/skills/verify-*/SKILL.md
Web (fetched and verified this run):
- https://arxiv.org/abs/2603.11027 — Beyond the Illusion of Consensus (Song, Zheng, Xu; 2026-03-11)
- https://arxiv.org/abs/2604.23178 — Judging the Judges (Soumik; 2026-04-25)
- https://arxiv.org/abs/2608.30373 — Beyond Consensus: Downward Bias and Role Asymmetry (Song, Kim, Eo, Park; 2026-08-31)
Web (search-summary level only, NOT fetched — treat as unverified):
- https://arxiv.org/pdf/2410.12784 — JudgeBench
- https://arxiv.org/pdf/2608.16003 — Prior Audit-Repair Context Shifts LLM Verifier Thresholds (title only)
- https://arxiv.org/pdf/2608.24419 — A Judge Should Know What Changed (title only)
- https://arxiv.org/pdf/2606.19544 — Reliability without Validity (title only)
- https://arxiv.org/pdf/2604.16790 — Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering (title only)
- https://orq.ai/blog/llm-juries-in-practice — practitioner framing on juries
- https://futureagi.com/blog/llm-as-judge-best-practices-2026/ — practitioner bias catalogue
Paywalls / access notes: none hit this run. Full PDFs of 2604.23178 and 2608.30373 were not retrieved; findings above come from their abstract pages only.