arXiv · 2609.15561
Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation
Abstract
Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuality-based metrics, their sensitivity and reliability remain underexplored. This paper introduces a meta-evaluation framework that systematically tests these metrics using controlled corruptions of gold standard answers. Our method generates ranked outputs with known degrees of degradation to probe how metrics capture nuanced changes in truthfulness. Our experiments reveal that pipeline-based methods, such as the RAGAS's factual correctness metric, better track degradation than LLM-as-judge approaches. We also propose a new variant of the factual correctness metric that provides a competitive and cost-efficient.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sarra Gharsallah, Adele Robaldo, Mariia Tokareva, Giovanni Gatti Pinheiro, Ilyana Guendouz, Raphaël Troncy, Paolo Papotti, Pietro Michiardi. 2026-09-14. Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation. https://arxiv.org/abs/2609.15561
Cite the original work for its findings. Save a collection to share your selection of sources.