arXiv · 2407.01122
Calibrated Large Language Models for Binary Question Answering
Abstract
Quantifying the uncertainty of predictions made by large language models (LLMs) in binary text classification tasks remains a challenge. Calibration, in the context of LLMs, refers to the alignment between the model's predicted probabilities and the actual correctness of its predictions. A well-calibrated model should produce probabilities that accurately reflect the likelihood of its predictions being correct. We propose a novel approach that utilizes the inductive Venn--Abers predictor (IVAP) to calibrate the probabilities associated with the output tokens corresponding to the binary labels. Our experiments on the BoolQ dataset using the Llama 2 model demonstrate that IVAP consistently outperforms the commonly used temperature scaling method for various label token choices, achieving well-calibrated probabilities while maintaining high predictive quality. Our findings contribute to the understanding of calibration techniques for LLMs and provide a practical solution for obtaining reliable uncertainty estimates in binary question answering tasks, enhancing the interpretability and trustworthiness of LLM predictions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Patrizio Giovannotti, Alexander Gammerman. 2024-07-01. Calibrated Large Language Models for Binary Question Answering. https://arxiv.org/abs/2407.01122
Cite the original work for its findings. Save a collection to share your selection of sources.