arXiv · 2609.32248
Anytime-Valid LLM Leaderboards via Benchmark-weighted and Block-Factorized e-Processes
Abstract
Large language model (LLM) leaderboards compare model capabilities by ranking models according to their mean performance on fixed benchmarks. However, variability in evaluation outcomes across runs may produce unsupported claims of model superiority on the benchmark, a risk compounded by leaderboard updates. In this paper, we propose BB-EDGE (Benchmark-Weighted and Block-Factorized e-processes for Directed Graph Evaluation), a principled framework that represents an LLM leaderboard as a directed graph whose edges certify pairwise mean-performance advantages, with anytime-valid family-wise error rate (FWER) control. Concretely, for each direction, BB-EDGE constructs an empirical-Bernstein e-process by factorizing evidence over protocol-defined blocks and assigning stakes proportional to the corresponding block weights, then applies direct e-Holm across these $e$-processes to certify directional advantages as edges. Theoretically, we characterize weight-proportional linear stakes under heterogeneous benchmark-average nulls and prove anytime FWER control under arbitrary within-block and cross-pair dependence. BB-EDGE further supports anytime-valid Top-$k$ certification and simultaneous rank intervals. Extensive experiments on synthetic data and four real-world benchmarks demonstrate that BB-EDGE maintains anytime FWER control while achieving high efficiency.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hongfu Gao, Songxin Zhang, Zejian Xie, Bingyi Jing, Zhou Wang, Yiming Liu. 2026-09-26. Anytime-Valid LLM Leaderboards via Benchmark-weighted and Block-Factorized e-Processes. https://arxiv.org/abs/2609.32248
Cite the original work for its findings. Save a collection to share your selection of sources.