arXiv · 2609.20058
Evaluating Explanation Methods by the Predictors They Induce
Abstract
Explanations of machine learning models are usually judged by criteria that are hard to compare. We propose a simpler test: if an explanation really describes how a model uses its features, it should be possible to rebuild the model's predictions from it. We turn each explanation into a predictor by reading each feature's effect and adding them up, and measure how well that predictor reproduces the model on unseen data. Nothing is fitted, so the score reflects the explanation itself. The test applies to any explanation that can be written as a function of the features; we demonstrate it on partial dependence plots (PDP), accumulated local effects (ALE), SHAP and LIME. We prove that summing partial dependence curves gives the best possible additive summary of a model when its features are independent, and that this fails when they are dependent. Across 13 real datasets and 9 synthetic designs and four model families, which method scores best depends entirely on feature dependence: where features are independent SHAP is slightly worse than PDP, exactly as the theory predicts; on dependent real data SHAP leads. Some widely used quality metrics even prefer a damaged explanation to an intact one.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jacob Selbæk, Hugo L. Hammer. 2026-09-17. Evaluating Explanation Methods by the Predictors They Induce. https://arxiv.org/abs/2609.20058
Cite the original work for its findings. Save a collection to share your selection of sources.