arXiv · 2604.11662
A Robust Evaluation of Probe Robustness: Lessons for Reliable OOD Uncertainty Quantification
Abstract
Recent work has shown that the hidden states of large language models contain signals useful for uncertainty estimation, motivating a growing interest in efficient probe-based approaches. Yet it remains unclear how robust existing methods are, with prior work reporting conflicting conclusions under substantially different evaluation settings. We address this by introducing ProbeDrift, a systematic evaluation framework for supervised uncertainty probes covering a wide range of OOD settings across models, tasks, and distributional shifts. Using ProbeDrift, we train over 2,000 probes to disentangle the effect of key design choices, showing poor robustness of current methods beyond near-OOD settings. We find that robustness is driven by design decisions that have a largely invisible effect in-distribution, including the choice of feature type, aggregation strategy, and training signal. We argue that robust uncertainty estimation requires robust evaluation. To support this, we release ProbeDrift as a lightweight Python library that contains the train and test splits underpinning our extensive evaluation. We also show how insights from our evaluation can directly lead to more robust methods through a simple Hybrid Back-Off (HBO) strategy.
Explore related subjects
Keep this discovery
Joe Stacey, Hadas Orgad, Kentaro Inui, Benjamin Heinzerling, Nafise Sadat Moosavi. 2026-08-30. A Robust Evaluation of Probe Robustness: Lessons for Reliable OOD Uncertainty Quantification. https://arxiv.org/abs/2604.11662
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.