06-reference/research

rereview critic diff visibility vs zero context

2026-09-29·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
llm-as-judgecritic-layerre-reviewconstruct-validityfresh-eyesverification

The conflict does not exist: one paper never asks the judge to see a diff, and the other never tests a fresh instantiation

The question

Verbatim: "Resolve the direct conflict between arXiv 2608.24419 ('A Judge Should Know What Changed' — construct validity requires the judge to see the diff) and arXiv 2608.16003 ('Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency'): should an RDCO re-review critic be shown the prior verdict and repair diff, or instantiated zero-context — and does the leniency shift survive a fresh instantiation or only a same-session one?"

Context: [[2026-09-25-llm-as-judge-surface-heuristic-bias]] proposed writing a zero-context RE-review rule into the verify-* SKILL.md family, then flagged 2608.24419 as a possible contradiction on the strength of its title alone. Neither paper had been dereferenced. This brief reads both in full and reports that the contradiction was an artifact of the title.

Citation verification (done first, as required)

Both IDs resolve, both titles match the queue's quotation. No correction needed.

ID Resolves to Authors Submitted
2608.24419 "A Judge Should Know What Changed: Construct Validity for LLM-as-a-Judge Evaluation" Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong 2026-08-25
2608.16003 "Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency" Parsa Mazaheri, Kasra Mazaheri 2026-08-17

Verified this run by dereferencing the arXiv metadata API and then pulling the full LaTeXML HTML of each paper (all sections plus appendices). The 2026-09-25 brief's "title only, no claim made about contents" caveat was correct to file, and the caution was warranted: the parenthetical gloss carried in the queue question, "construct validity requires the judge to see the diff," is not a claim 2608.24419 makes. That gloss was inferred from the title by the surfacing run, not read from the paper.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

Recommendation, awaiting the founder's read: keep the re-review critic zero-context, write the rule down explicitly, and cite 2603.12123 for it rather than 2608.16003. The zero-context rule survives contact with 2608.24419 untouched, because that paper judges pointwise and never proposes a diff-visible judge. The rule should not, however, be justified by 2608.16003, and the 09-25 brief's proposed justification ("a judge that has just watched a repair is primed to accept it") overreaches the evidence in four separate ways: the episode is about a different item, the delivery is prepended turns rather than a handoff, no frontier model was tested and one screened model reversed, and the shift was net beneficial at the measured operating point. Cross-Context Review is the honest citation: it separates production from review, it includes a context-aware subagent arm that CCR beat by 4.8 F1 points at p=0.004, and its SR2 null rules out repetition as the explanation. Its weaknesses are small n and a single author; label it accordingly rather than laundering it.

FIRST review and RE-review get the same answer, for different reasons, and the brief should say so plainly. First review: zero-context, unchanged, supported by [[2026-05-19-verification-as-independent-worker-pattern]] and now by CCR's direct measurement. Re-review: zero-context as the default, but on the weaker ground that it is the status quo, it costs one subagent, and the direction of the alternative is uncontrolled rather than proven harmful. That is a materially more honest framing than "the literature says the diff-visible critic is lenient," and it is the framing a future /improve cycle can actually falsify once the verdict log exists.

The genuinely new build this brief argues for is a defect-closed check that lives in the harness, not in the judge. 2608.24419's architecture is the template and RDCO half-built it already: generation, verification, and judging assigned to disjoint model families, with the original-vs-edit pair given to a mechanical verifier while the judge stays pointwise. Translated, the re-review round should be an AND of two independent instruments. (a) A fresh zero-context critic re-scores the whole artifact from the rubric, exactly as today, seeing no prior verdict, no defect list, and no diff. (b) A separate narrow checker receives only the current artifact plus each prior defect stated as a standalone question ("does this note resolve every cited arXiv ID to a matching title?"), with no prior verdict, no repair diff, and no producer notes, and answers per defect. The harness holds the history and ANDs the results. This is the reconciliation the queue was looking for: the diff belongs to the harness, the rubric belongs to the judge, and "know what changed" is a property the loop should have, not a document the critic should read. verify-pdf-output's existing "escalate with verdict history + diff list" and station-critic's critic-feedback-iter-<N>.md are the two places RDCO already does this correctly; the gap is that both feed the producer, and nothing checks defect closure independently.

None of this is measurable, and the measurement design is now specific enough to build. 2608.16003's methodological warning transfers to RDCO verbatim and is the most immediately usable thing in either paper: "a metric computed on flagged items alone will read that as an improvement." If RDCO ever notices that round-2 critics flag less than round-1 critics, that number is uninterpretable on its own. It must be paired with a detection rate on artifacts carrying known-seeded defects, from the same run. That is the polished-wrong-twin experiment from [[2026-09-25-llm-as-judge-surface-heuristic-bias]] extended to a second round: seed a defect, let the producer repair a different defect, and count how many round-2 critics still catch the untouched one, under both a fresh critic and a context-loaded critic. 2606.15474's anchor-set design suggests the cheap standing version: hold a small frozen set of human-labeled artifacts and have the current critic re-score them on an interleave, so a leniency drift in RDCO's own critics is attributable rather than ambiguous. All of it is still downstream of the verdict log that [[2026-07-03-verify-skills-pass-disagreement-audit]] named on 2026-07-03 and that remains unbuilt.

Ray has no authority to edit any SKILL.md on this, and has not. The rule change proposed here is: add an explicit RE-REVIEW clause to the fresh-eyes pre-flight in verify-vault-write, verify-strategic-output, verify-dispatch and verify-pdf-output, stating that a re-review is a new subagent that receives the current artifact and the rubric only, and separately add the defect-closed checker as a second instrument. Both await the founder's read.

Why this is in the vault

This unblocks the specific SKILL.md edit that [[2026-09-25-llm-as-judge-surface-heuristic-bias]] deliberately parked: it establishes that 2608.24419 does not contradict a zero-context re-review rule, replaces the overreaching 2608.16003 justification with a citation that actually measures fresh-session review, and converts the question into a concrete two-instrument re-review design for the verify-* family. It also records a caught failure in RDCO's own research pipeline: /deep-research generated a "direct conflict" follow-up from a paper title it had never dereferenced, and the conflict evaporated on contact with the text.

Open follow-ups

Related

Sources

Vault:

Skill files inspected directly on disk 2026-09-29 (not inferred):

Web (FULL TEXT fetched and read this run, via arXiv LaTeXML HTML):

Web (ABSTRACT-LEVEL ONLY, verified via the arXiv metadata API — treat claims as unconfirmed beyond the abstract):

Not fetched, carried from 2026-09-25 as unverified:

Paywalls / access notes: none. http://export.arxiv.org 301-redirects; the HTTPS endpoint with -L works. Research caps respected: 1 QMD query, 1 WebSearch, 2 full-text fetches plus 2 metadata lookups, 5 vault docs and 5 skill files read.