arXiv · 2610.09820
The Impact of Backbone Evolution on LLM-Based Relevance Assessments
Abstract
LLMs are evolving rapidly, with newer models offering stronger capabilities. This suggests that in LLM-based relevance judging, more capable models will achieve higher agreement with human judgements under the same prompt. We challenge this understanding by investigating the behavior of LLM-based relevance judges under backbone evolution. Keeping the prompts fixed, we evaluate a representative single-prompt (UMBRELA) and a rubric-based prompt (EXAM) across sequential model versions of commercial (Gemini, GPT) and open-weight (Qwen, Llama) models. Overall, we find no consistent evidence that newer versions lead to better relevance judges. Crucially, similar or improved aggregate performance does not imply judgment stability: correct judgements made by an earlier version of an LLM backbone are not necessarily preserved by later versions. We investigate the potential drivers of these regressions. Our findings caution against the assumption that judging prompts designed and validated for one backbone version will perform equivalently or better when the model is updated, even within the same family.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chuting Yu, Guido Zuccon, Teerapong Leelanupab. 2026-10-07. The Impact of Backbone Evolution on LLM-Based Relevance Assessments. https://arxiv.org/abs/2610.09820
Cite the original work for its findings. Save a collection to share your selection of sources.