Pinned and Still Unstable: Within-Judge Verdict Variance and the Noise Floor of LLM-as-Judge Leaderboards
Modern LLM evaluation rests on the assumption that pinning a judge to a fixed model snapshot and decoding at temperature zero yields reproducible verdicts. The paper shows this assumption fails, and argues it is a property of how LLM-as-Judge is operationalized on cloud serving infrastructure rather than of any particular model family.
The study covers four frontier judges, all served through a single major enterprise cloud platform, across three standard benchmarks: Arena-Hard, AlpacaEval 2, and MT-Bench. Identical inputs to the same pinned, temperature-zero judge produce different verdicts across re-runs, with per-item flip rates of roughly 5% on average and about 40% on the close-call items that decide leaderboard margins. The per-judge magnitude spans a 40x range, from 0.13% to nearly 10%.
To characterize this, the authors introduce metrics tailored to the instability: per-item flip rate, a two-part stability profile (waver fraction and conditional intensity), and adjacency separability. They then report what the variance does and does not do to rankings. For a single judge, the aggregate ranking is stable — 0% top-K instability and 0% pooled winner flip — but precision degrades: under a paired hierarchical bootstrap, roughly one-fifth to three-quarters of adjacent leaderboard positions are statistically indistinguishable, a noise floor the authors attribute primarily to finite prompt sampling rather than to the judge itself.
Across judges, leaderboards agree on the coarse ordering but diverge in the middle, with Kendall's tau as low as 0.42-0.64 between families on Arena-Hard. Of 13 published head-to-head ranking claims the authors re-judge, 5 fail under a defensible judge swap or re-run. They argue that leaderboards report unhedged point estimates that misrepresent the noise floor of the instrument, and propose a minimal, low-cost reporting protocol: several judge re-runs, published stability profiles and adjacency intervals, and results under at least two judges from different families.