The conflict does not exist: one paper never asks the judge to see a diff, and the other never tests a fresh instantiation
The question
Verbatim: "Resolve the direct conflict between arXiv 2608.24419 ('A Judge Should Know What Changed' — construct validity requires the judge to see the diff) and arXiv 2608.16003 ('Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency'): should an RDCO re-review critic be shown the prior verdict and repair diff, or instantiated zero-context — and does the leniency shift survive a fresh instantiation or only a same-session one?"
Context: [[2026-09-25-llm-as-judge-surface-heuristic-bias]] proposed writing a zero-context RE-review rule into the verify-* SKILL.md family, then flagged 2608.24419 as a possible contradiction on the strength of its title alone. Neither paper had been dereferenced. This brief reads both in full and reports that the contradiction was an artifact of the title.
Citation verification (done first, as required)
Both IDs resolve, both titles match the queue's quotation. No correction needed.
| ID | Resolves to | Authors | Submitted |
|---|---|---|---|
| 2608.24419 | "A Judge Should Know What Changed: Construct Validity for LLM-as-a-Judge Evaluation" | Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong | 2026-08-25 |
| 2608.16003 | "Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency" | Parsa Mazaheri, Kasra Mazaheri | 2026-08-17 |
Verified this run by dereferencing the arXiv metadata API and then pulling the full LaTeXML HTML of each paper (all sections plus appendices). The 2026-09-25 brief's "title only, no claim made about contents" caveat was correct to file, and the caution was warranted: the parenthetical gloss carried in the queue question, "construct validity requires the judge to see the diff," is not a claim 2608.24419 makes. That gloss was inferred from the title by the surfacing run, not read from the paper.
What we already know (from the vault)
- The zero-context rule is already RDCO's operating doctrine, and it is already in the skills. [[2026-05-19-verification-as-independent-worker-pattern]] makes independence structural rather than motivational.
verify-vault-write/SKILL.mdgoes further with a pre-flight the critic runs on itself: "I am a fresh-eyes sub-agent with zero build context. If I have prior context on how this note was written, ABORT and surface to parent." The proposal in the 09-25 brief was therefore never a new rule; it was making the re-review case explicit in a rule that currently reads as if it only governs first review. - The harness already holds verdict history without giving it to the judge, in exactly one place.
verify-pdf-output/SKILL.mdcaps iteration at 3 cycles and escalates "with the PDF + critic verdict history + diff list," and escalates immediately "if two consecutive verdicts return the SAME failing checks." That history lives in the parent, not in the critic's context.station-criticdoes the same thing withcritic-feedback-iter-<N>.md, which the next code-author iteration reads, not the next critic. This split is the answer to the question, and RDCO built it before asking. - Agreement between RDCO critics carries almost no information. [[2026-07-30-multi-llm-critic-ensembles-differentiated-rubrics]] found zero unique catches from a second vendor, and [[2026-09-25-llm-as-judge-surface-heuristic-bias]] tied that to the 62%-rubric-echo result. Relevant here because a re-review critic that shares a rubric AND a prior verdict with its predecessor is the maximally correlated configuration available.
- The blind spot this brief has to respect: there is still no verdict log. [[2026-07-03-verify-skills-pass-disagreement-audit]] named it the binding constraint; the 09-25 brief re-verified it absent on disk twelve weeks later. Any leniency claim about RDCO's own loop is unfalsifiable today, including the one recommended below.
- RDCO has already noticed that critic feedback returns as prose, not as machine-checkable failures. [[2026-07-28-amazon-engineer-agentic-signal-loop-not-reading-the-diff]] drew the distinction; [[2026-09-15-every-typesafe-jev-vibe-check]] observed that every RDCO critic is a one-shot post-hoc gate. Both bear directly on the "defect-closed check" proposed below.
What the web says
- 2608.24419's title is rhetorical, and the diff is never an input to the judge under test. The judge is formalised pointwise: one item plus one criterion in, one verdict out. "Know what changed" means the judge's verdict should move when the construct moves. The original-vs-edit pair is shown only to (a) a mechanical verifier in a disjoint model family and (b) human annotators in randomised order with provenance withheld. Judging is assigned to "family C, never A or B." The paper's own "paired mode" is a RewardBench/MT-Bench A-vs-B preference between two independent candidate answers, explicitly not original-vs-edit. (https://arxiv.org/abs/2608.24419)
- 2608.24419's actual recommendation is a reporting standard, not an input change. Its §6.1 "what a judge author can do on Monday" asks for label provenance, the best impoverished predictor's agreement, a human ceiling with an interval, a label-permutation null, probing in both directions, and a released provenance field. Nowhere does it recommend giving a judge a diff, a prior version, or a revision history. It says nothing at all about re-review, repair loops, or iterative refinement; greps for those terms return only unrelated hits.
- What 2608.24419 does establish is far more useful to RDCO than the supposed conflict. At matched invariance S = 0.945, mean construct sensitivity across 7 judges is R = 0.319. Judges are 94.5% stable under surface edits and change their verdict only ~32% of the time when the underlying truth changes. Sensitivity splits by edit type: R_scope = 0.383 vs R_strength = 0.262, a +0.121 gap with the same sign on all 7 judges. Per-slot flip rates put hedging removal lowest at 0.210 and intensifier addition at 0.235. Per-judge R ranges from claude-haiku-4.5 at 0.561 down to llama-4-scout at 0.093. A surface-only predictor reproduced 67.4% of MT-Bench human votes in paired mode.
- 2608.16003's manipulation is a prior audit-repair episode about a DIFFERENT item, prepended as turns in the same context window. Disjointness is load-bearing and enforced by a 929/50/122 split of ProcessBench clean traces; the target request is byte-identical across conditions; the control is a length-matched filler turn in which "nothing in the request or the reply names correctness, error, review or repair." The false-alarm rate drops in 15 of 15 model-by-wording combinations, 2.8 to 11.5 points, 9-25% relative. Signal-detection analysis puts the move in the criterion (survives Holm-Bonferroni in 13 of 15) and not in d' (survives in 0 of 15, though SE(Δd') is exactly twice SE(Δc) by design, so the d' test is half as sensitive). (https://arxiv.org/abs/2608.16003)
- 2608.16003 does not test fresh instantiation anywhere. There is no condition in which a verdict is summarised and handed to a clean model call. Every measured condition is a single context with the episode as literal prepended turns. The nearest things are an Appendix B "prospective ladder" of stated future repair obligations inside the same prompt (which moves d' downward in 13 of 13 surviving contrasts, so it is not free either) and relabelled self-vs-peer attribution cells, which move nothing. This is a direct answer to the queue's second sub-question: the literature does not separate same-session from fresh-instantiation leniency, because it has only ever measured the same-session case.
- The leniency was net beneficial at their operating point, which breaks the intuitive reading. A hand audit of 50 baseline false alarms found 41 (82%) "simply wrong"; arithmetic and algebraic flags account for 97% of the false-alarm reduction, i.e. the flags removed were the fabricated ones. Detection on a matched 929-trace incorrect arm fell only 0.3 to 1.6 points against 2.8 to 6.5 points of false-alarm reduction, and balanced accuracy rose in 15 of 15 combinations. The authors' warning is about control, not direction: "the threshold moved without anyone asking it to, and whether a threshold move helps depends on an operating point a pipeline change can silently alter."
- A third paper, not in the queue, actually answers the fresh-instantiation question affirmatively. "Cross-Context Review: Improving LLM Output Quality by Separating Production and Review Sessions" (Tae-Eun Song, arXiv 2603.12123, 2026-03-12) ran 30 artifacts with 150 injected errors across 360 reviews under four conditions. Cross-Context Review in a fresh session with no production history reached F1 28.6%, beating same-session self-review at 24.6% (p=0.008, d=0.52), repeated same-session self-review at 21.7% (p<0.001, d=0.72), and context-aware subagent review at 23.8% (p=0.004, d=0.57). The interpretive lever: reviewing twice in the same session did not beat reviewing once (p=0.11), which rules out repetition and pins the benefit on context separation itself. Abstract-level only, not full text; single author, small n, and all four F1 values are low in absolute terms. (https://arxiv.org/abs/2603.12123)
- A fourth, for the instrumentation gap: "Who Drifted: the System or the Judge?" (Yitao Li, arXiv 2606.15474) uses a fixed human-labeled anchor set that the current judge re-scores at a steady interleave, plus a betting e-process, to return a verdict in {none, system, judge}. It detected a silent judge version bump as judge drift in 60/60 runs with zero misattribution, against an industry-default rolling z-test that false-alarmed on 75% of drift-free streams. Abstract-level only. (https://arxiv.org/abs/2606.15474)
Convergences and contradictions
- The premise of the question is false, and that is the headline. 2608.24419 and 2608.16003 do not touch. One is a pointwise one-shot measurement paper whose "diff" is an experimenter's instrument applied across disjoint model families; the other is a context-priming study on a different-item episode. Neither addresses re-review of the same artifact after a repair. The queue's "direct conflict" was generated by reading a rhetorical title as a design prescription, which is the exact failure mode [[2026-07-19-constitution-or-collapse-citation-verification]] caught in RDCO's own corpus, reproduced here in RDCO's own research pipeline.
- The real contradiction is between the two papers' predictions for a re-review loop, and it runs in opposite directions. 2608.16003 predicts a context-loaded critic drifts toward PASS. 2608.24419 predicts a critic of any kind fails to notice that the construct changed 68% of the time, and is worst precisely at strength edits: hedging removal flips 0.210, intensifier addition 0.235. Adding a hedge is itself classified as a strength edit in that paper (hedge padding was cut from the control set for that reason). The single most common RDCO repair is adding hedges in response to an overconfidence FAIL from
verify-strategic-output. So the leniency literature says the re-review rubber-stamps and the construct-validity literature says the re-review is roughly verdict-blind to whether the repair happened at all. Those are two different failure modes of the same loop and neither is fixed by showing the judge a diff. - Transfer to RDCO is unestablished on both legs, and one pre-screened model reversed. 2608.16003 audited three open-weight models served locally (Qwen3.6-27B, Qwen3.6-35B-A3B, Ministral-3-14B) and explicitly screened out Gemma-4-31B because its effect was significant and reversed at +1.58pp; the authors scope the direction to "framing-responsive verifiers rather than every verifier we tried." No frontier or closed judge was tested. RDCO's critics are Claude-family. 2608.24419 did test claude-haiku-4.5, and it had the highest construct sensitivity of the seven at R = 0.561, which is both reassuring and still barely one flip in two.
- Convergence with the 09-25 brief on surface bias. The surface-only predictor recovering 67.4% of MT-Bench human votes is an independent replication of that brief's core finding, from a different research group and a different instrument.
Synthesis for RDCO
Recommendation, awaiting the founder's read: keep the re-review critic zero-context, write the rule down explicitly, and cite 2603.12123 for it rather than 2608.16003. The zero-context rule survives contact with 2608.24419 untouched, because that paper judges pointwise and never proposes a diff-visible judge. The rule should not, however, be justified by 2608.16003, and the 09-25 brief's proposed justification ("a judge that has just watched a repair is primed to accept it") overreaches the evidence in four separate ways: the episode is about a different item, the delivery is prepended turns rather than a handoff, no frontier model was tested and one screened model reversed, and the shift was net beneficial at the measured operating point. Cross-Context Review is the honest citation: it separates production from review, it includes a context-aware subagent arm that CCR beat by 4.8 F1 points at p=0.004, and its SR2 null rules out repetition as the explanation. Its weaknesses are small n and a single author; label it accordingly rather than laundering it.
FIRST review and RE-review get the same answer, for different reasons, and the brief should say so plainly. First review: zero-context, unchanged, supported by [[2026-05-19-verification-as-independent-worker-pattern]] and now by CCR's direct measurement. Re-review: zero-context as the default, but on the weaker ground that it is the status quo, it costs one subagent, and the direction of the alternative is uncontrolled rather than proven harmful. That is a materially more honest framing than "the literature says the diff-visible critic is lenient," and it is the framing a future /improve cycle can actually falsify once the verdict log exists.
The genuinely new build this brief argues for is a defect-closed check that lives in the harness, not in the judge. 2608.24419's architecture is the template and RDCO half-built it already: generation, verification, and judging assigned to disjoint model families, with the original-vs-edit pair given to a mechanical verifier while the judge stays pointwise. Translated, the re-review round should be an AND of two independent instruments. (a) A fresh zero-context critic re-scores the whole artifact from the rubric, exactly as today, seeing no prior verdict, no defect list, and no diff. (b) A separate narrow checker receives only the current artifact plus each prior defect stated as a standalone question ("does this note resolve every cited arXiv ID to a matching title?"), with no prior verdict, no repair diff, and no producer notes, and answers per defect. The harness holds the history and ANDs the results. This is the reconciliation the queue was looking for: the diff belongs to the harness, the rubric belongs to the judge, and "know what changed" is a property the loop should have, not a document the critic should read. verify-pdf-output's existing "escalate with verdict history + diff list" and station-critic's critic-feedback-iter-<N>.md are the two places RDCO already does this correctly; the gap is that both feed the producer, and nothing checks defect closure independently.
None of this is measurable, and the measurement design is now specific enough to build. 2608.16003's methodological warning transfers to RDCO verbatim and is the most immediately usable thing in either paper: "a metric computed on flagged items alone will read that as an improvement." If RDCO ever notices that round-2 critics flag less than round-1 critics, that number is uninterpretable on its own. It must be paired with a detection rate on artifacts carrying known-seeded defects, from the same run. That is the polished-wrong-twin experiment from [[2026-09-25-llm-as-judge-surface-heuristic-bias]] extended to a second round: seed a defect, let the producer repair a different defect, and count how many round-2 critics still catch the untouched one, under both a fresh critic and a context-loaded critic. 2606.15474's anchor-set design suggests the cheap standing version: hold a small frozen set of human-labeled artifacts and have the current critic re-score them on an interleave, so a leniency drift in RDCO's own critics is attributable rather than ambiguous. All of it is still downstream of the verdict log that [[2026-07-03-verify-skills-pass-disagreement-audit]] named on 2026-07-03 and that remains unbuilt.
Ray has no authority to edit any SKILL.md on this, and has not. The rule change proposed here is: add an explicit RE-REVIEW clause to the fresh-eyes pre-flight in verify-vault-write, verify-strategic-output, verify-dispatch and verify-pdf-output, stating that a re-review is a new subagent that receives the current artifact and the rubric only, and separately add the defect-closed checker as a second instrument. Both await the founder's read.
Why this is in the vault
This unblocks the specific SKILL.md edit that [[2026-09-25-llm-as-judge-surface-heuristic-bias]] deliberately parked: it establishes that 2608.24419 does not contradict a zero-context re-review rule, replaces the overreaching 2608.16003 justification with a citation that actually measures fresh-session review, and converts the question into a concrete two-instrument re-review design for the verify-* family. It also records a caught failure in RDCO's own research pipeline: /deep-research generated a "direct conflict" follow-up from a paper title it had never dereferenced, and the conflict evaporated on contact with the text.
Open follow-ups
- Does the 2608.16003 leniency shift reproduce when the prior episode concerns the same artifact under review, rather than a held-out different item? The paper's disjointness constraint is load-bearing and deliberately excludes RDCO's actual configuration, so the closest-to-RDCO condition is unmeasured in the literature.
- Does the shift reproduce on a frontier closed judge? All three audited models were open-weight 14-35B served locally, and the one screened-out Gemma model reversed direction, so the authors scope the finding to "framing-responsive verifiers." No published measurement on a Claude-family verifier was found this run.
- What is construct sensitivity R for a rubric-gated pointwise PASS/FAIL, rather than for a graded verdict thresholded by the experimenters? 2608.24419's Corollary 2 states the two modes are not orderable and that "a third mode's VP is not deducible from ours," so RDCO's gate mode is out of scope for its own numbers.
- Is the R_scope > R_strength gap stable for edits that add hedges rather than remove them? The paper classifies hedge padding as a strength edit but excluded it from the control set, and every measured strength slot moves commitment upward. The repair direction RDCO actually produces is the untested one.
- Has anyone measured whether a defect-stated-as-standalone-question checker outperforms a diff-visible checker at confirming a repair landed? This is the empirical core of the proposed design and no source found this run tests it either way.
- Does the Cross-Context Review result hold at RDCO's artifact class and length? CCR used 30 artifacts of code, technical documents and presentation scripts with injected errors; F1 was 28.6% in absolute terms, which is low enough that the ranking may not transfer to 3,000-word research briefs.
- Enumerate which of the nine debiasing strategies in arXiv 2604.23178 produced which deltas. Carried unresolved from 2026-09-25; the abstract withheld the mapping and the full PDF is still unfetched.
Related
- [[2026-09-25-llm-as-judge-surface-heuristic-bias]]
- [[2026-07-03-verify-skills-pass-disagreement-audit]]
- [[2026-07-30-multi-llm-critic-ensembles-differentiated-rubrics]]
- [[2026-05-19-verification-as-independent-worker-pattern]]
- [[2026-05-20-verify-stack-two-gate-pass-fail-architecture]]
- [[2026-07-19-constitution-or-collapse-citation-verification]]
- [[2026-07-26-autoreview-skill-teardown-second-model-critic-design]]
- [[2026-07-28-amazon-engineer-agentic-signal-loop-not-reading-the-diff]]
- [[2026-09-15-every-typesafe-jev-vibe-check]]
- [[2026-05-23-improve-cron-design-spec]]
- [[2026-05-19-cai-critic-graduation-per-axis-threshold]]
Sources
Vault:
- ~/rdco-vault/06-reference/research/2026-09-25-llm-as-judge-surface-heuristic-bias.md
- ~/rdco-vault/06-reference/research/2026-07-03-verify-skills-pass-disagreement-audit.md
- ~/rdco-vault/06-reference/research/2026-07-30-multi-llm-critic-ensembles-differentiated-rubrics.md
- ~/rdco-vault/02-sops/2026-05-19-verification-as-independent-worker-pattern.md
- ~/rdco-vault/02-sops/2026-05-20-verify-stack-two-gate-pass-fail-architecture.md
- ~/rdco-vault/06-reference/research/2026-07-19-constitution-or-collapse-citation-verification.md
- ~/rdco-vault/06-reference/2026-07-26-autoreview-skill-teardown-second-model-critic-design.md
- ~/rdco-vault/06-reference/2026-07-28-amazon-engineer-agentic-signal-loop-not-reading-the-diff.md
- ~/rdco-vault/06-reference/2026-09-15-every-typesafe-jev-vibe-check.md
- ~/rdco-vault/06-reference/research/2026-05-23-improve-cron-design-spec.md
- ~/rdco-vault/06-reference/research/2026-05-19-cai-critic-graduation-per-axis-threshold.md
Skill files inspected directly on disk 2026-09-29 (not inferred):
- ~/.claude/skills/verify-vault-write/SKILL.md (fresh-eyes pre-flight; "re-invoke" at line 223)
- ~/.claude/skills/verify-strategic-output/SKILL.md (line 218, max 2 iterate rounds)
- ~/.claude/skills/verify-dispatch/SKILL.md (line 226)
- ~/.claude/skills/verify-pdf-output/SKILL.md (line 227, verdict history held by the parent)
- ~/.claude/skills/station-critic/SKILL.md (critic-feedback-iter-N.md feeds the code-author, not the next critic)
Web (FULL TEXT fetched and read this run, via arXiv LaTeXML HTML):
- https://arxiv.org/abs/2608.24419 — A Judge Should Know What Changed: Construct Validity for LLM-as-a-Judge Evaluation (Chen, Chen, Lin, Vong; 2026-08-25). All sections plus Appendices A-H.
- https://arxiv.org/abs/2608.16003 — Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency (Mazaheri, Mazaheri; 2026-08-17). Full text; code repo cited at https://github.com/parsa-mz/crtitxer (not inspected).
Web (ABSTRACT-LEVEL ONLY, verified via the arXiv metadata API — treat claims as unconfirmed beyond the abstract):
- https://arxiv.org/abs/2603.12123 — Cross-Context Review (Tae-Eun Song; 2026-03-12)
- https://arxiv.org/abs/2606.15474 — Who Drifted: the System or the Judge? (Yitao Li; 2026-06-13)
Not fetched, carried from 2026-09-25 as unverified:
- https://arxiv.org/abs/2604.23178 — Judging the Judges (nine-strategy mapping still unresolved)
Paywalls / access notes: none. http://export.arxiv.org 301-redirects; the HTTPS endpoint with -L works. Research caps respected: 1 QMD query, 1 WebSearch, 2 full-text fetches plus 2 metadata lookups, 5 vault docs and 5 skill files read.