arXiv · 2609.33044
Pinned and Still Unstable: Within-Judge Verdict Variance and the Noise Floor of LLM-as-Judge Leaderboards
Abstract
Modern LLM evaluation assumes that pinning a judge to a fixed model snapshot and decoding at temperature zero yields reproducible verdicts. We show this assumption fails as a property of how LLM-as-Judge is operationalized on cloud serving infrastructure, not of any particular model family. Across four frontier judges all served via a single major enterprise cloud platform and three standard benchmarks (Arena-Hard, AlpacaEval 2, MT-Bench), identical inputs to the same pinned, temperature-zero judge produce different verdicts across re-runs: per-item flip rates of roughly 5% on average and about 40% on the close-call items that decide leaderboard margins, with a per-judge magnitude spanning a 40x range (from 0.13% to nearly 10%). We introduce metrics tailored to this instability: per-item flip rate, a two-part stability profile (waver fraction and conditional intensity), and adjacency separability - and report what the variance does and does not do to rankings. For a single judge the aggregate ranking is stable (0% top-K instability, 0% pooled winner flip); what degrades is precision: under a paired hierarchical bootstrap, roughly one-fifth to three-quarters of adjacent leaderboard positions are statistically indistinguishable, a noise floor we attribute primarily to finite prompt sampling rather than to the judge. Across judges, leaderboards agree on the coarse ordering but diverge in the middle (Kendall's tau as low as 0.42-0.64 between families on Arena-Hard), and of 13 published head-to-head ranking claims we re-judge, 5 fail under a defensible judge swap or re-run. We argue that leaderboards report unhedged point estimates that misrepresent the noise floor of the instrument, and we propose a minimal, low-cost reporting protocol: several judge re-runs, published stability profiles and adjacency intervals, and results under at least two judges from different families.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Krishna Chytanya Ayyagari. 2026-09-27. Pinned and Still Unstable: Within-Judge Verdict Variance and the Noise Floor of LLM-as-Judge Leaderboards. https://arxiv.org/abs/2609.33044
Cite the original work for its findings. Save a collection to share your selection of sources.