HalluTracer: Pre-Decoding Truthfulness Prediction via Depth-Averaged Probe-Logit
Internal-state probes enable truthfulness prediction before a large language model generates an answer. When detectors change both the layers they read and the rules used to combine them, the source of improved prediction becomes difficult to identify. We separate these choices and find that retaining more layers improves prediction even under fixed equal weighting. An exact Fisher-ratio decomposition explains why the additional benefit of linear reweighting is limited on these probe scores: information inlayer-wise differences largely overlaps with that captured by the depth mean. Estimating additional weights can then offset this small benefit when the data used to fit them are limited. These findings motivate our proposed method HalluTracer, which averages layer-wise probe logits to predict truthfulness before decoding. Across six models and four benchmarks, including TruthfulQA, HalluTracer achieves the highest area under the receiver operating characteristic curve (AUROC) in 23 of 24 model--benchmark pairs among the compared methods. The results support using evidence from across the network without requiring a correspondingly more flexible aggregation rule, clarifying the distinct roles of layer selection and weighting in pre-decoding detection.