TY - RPRT TI - When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation AU - Wenda Xu AU - Sweta Agrawal AU - Vilém Zouhar AU - Markus Freitag AU - Daniel Deutsch PY - 2026 UR - https://arxiv.org/abs/2509.26600 ID - 2509.26600 ER -