arXiv · 2509.25671
The Flaw of Averages: Measuring Benchmark-Level Distributional Robustness
Abstract
Benchmarks are central to measuring progress in language models, but aggregate scores can obscure substantial variation across subdomains, making models appear broadly competent despite concentrated strengths and weaknesses. We study this issue as benchmark-level distributional robustness: whether aggregate scores faithfully reflect performance across benchmark subdomains. We operationalize this notion with benchmark Harmony, an entropy-based measure of how uniformly model performance is distributed across subdomains. Measuring Harmony on 19 language model benchmarks across five model families, we find substantial variation in benchmark-level distributional robustness. Low-Harmony benchmarks are more likely to yield aggregate scores that overstate broad competence, whereas high-Harmony benchmarks provide more representative summaries of model capability. Rebalancing benchmarks by pruning overrepresented subdomains to increase Harmony substantially shifts aggregate scores for low-Harmony benchmarks, but leaves high-Harmony benchmarks comparatively stable. For example, while BoolQ remains comparatively stable as Harmony increases, PubMedQA, which evaluates performance in a medically consequential domain, exhibits substantial, often statistically significant, shifts in aggregate accuracy. Together, these findings show that aggregate scores can misrepresent broad competence when performance is unevenly distributed. We therefore recommend reporting benchmark Harmony alongside aggregate accuracy as a diagnostic of benchmark representativeness when interpreting claims about broad model competence.
Explore related subjects
Keep this discovery
Arda Uzunoglu, Tianjian Li, Daniel Khashabi. 2025-09-30. The Flaw of Averages: Measuring Benchmark-Level Distributional Robustness. https://arxiv.org/abs/2509.25671
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.