arXiv Science⌕ Search

arXiv · 2610.08377

Data geometry preserves prediction but reshapes explanation

Abstract

Across scientific domains, empirical data often give rise to positive semidefinite matrices that encode similarities, couplings or interactions and induce natural geometries. We investigate whether changing the geometry used to compare the same empirical representations preserves subject-level performance and the local structures from which explanations are derived, a question we term the prediction explanation invariance problem. We address this problem across morphometric MRI, fMRI and EEG data by comparing Frobenius and trace geometries applied to the same empirical matrices. Predictive performance was broadly preserved across geometries, but comparable performance did not imply agreement in subject-level decision scores or predicted labels. Predictive similarity also masked geometry-dependent differences in local neighbourhoods, perturbation sensitivities and explanatory rankings. Dimensionality reduction made the two subject-space geometries progressively more concordant, while classification performance and decision-level agreement declined. These results identify data geometry as a hidden degree of freedom in explainability: similar performance does not guarantee invariant decisions or explanations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nicola Amoroso, Mario Caruso, Marianna La Rocca, Loredana Bellantuono, Tommaso Maggipinto, Michele Morelli, Sabina Tangaro, Marco Tatullo, Roberto Bellotti, Ester Pantaleo, Alfonso Monaco. 2026-10-06. Data geometry preserves prediction but reshapes explanation. https://arxiv.org/abs/2610.08377

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Modular Aggregation as a Debiasing Method for Non-Stationary Discrete Sources: Convergence and Numerical Validatio

We analyze \emph{modular aggregation}---summing $N$ independent outcomes modulo $m$---as a post-processing method for extracting nearly uniform randomness from biased discrete sources. Using discrete Fourier analysis over the cyclic group $\mathbb{Z}_m$, we prove exponential convergence of the output distribution to uniformity, with a rate determined by the largest non-trivial Fourier modulus. The result applies to independent non-stationary (non-IID) sources under a uniform spectral-gap condition on the non-trivial Fourier modes. Numerical simulations under several bias regimes, including cyclic drift and extreme cyclic bias, are used as finite-sample diagnostics and illustrate the theoretical predictions in comparison with Peres extraction and SHA-256 post-processing. The robustness of modular aggregation comes at a retention cost of order $1/N$, yielding an explicit trade-off between statistical quality and throughput.

physics.data-an↗

Developing Multi-Dimensional, Sequential sWeighting Routines for Reaction Selection with Emphasis on Kaon Identification at CLAS12

In nuclear and particle physics experiments, event selection is often a multidimensional classification problem which involves several correlated observables. The widely used sWeight technique provides a statistically rigorous means of separating signal and background using discriminating variables, but its conventional formulation is generally applied to a single discriminating variable or a simultaneous multidimensional fit. In this paper we present a novel extension of the widely used sWeight technique, in which an arbitrary number of discriminating variables can be used to construct a multidimensional sequential framework. This method allows the extraction of signal distributions in controlled variables in a rigorous way. Particular attention is given to the treatment of signal and background correlations, as well as the propagation of statistical uncertainties through successive weighting stages. The method yields high-purity signal samples even in the presence of large and correlated backgrounds. As a benchmark application, the validity of this technique is demonstrated on kaon identification in a dataset collected with the CLAS12 detector at the Thomas Jefferson National Laboratory for cascade baryon searches. The final workflow is shown to be robust and consistent across different run periods and beam/detector conditions, demonstrating minimal systematical uncertainties associated with the method. Statistical uncertainties are proved to be evaluable using a bootstrap-based procedure. The method provides a general framework for multidimensional event weighting and is intended for application in complex analyses requiring robust signal selection in the presence of complex correlated backgrounds.

physics.data-an↗

Statistical validation of calorimeter inpainting with generative diffusion priors

Localized detector inefficiencies produce incomplete calorimeter data that limit the ability to perform precision measurements. We address this problem in relativistic heavy-ion collisions from a Bayesian perspective using pretrained calorimeter diffusion models as priors to reconstruct the missing signal conditioned on surrounding measurements. In this work, we conduct a systematic comparison of several diffusion-based inpainting algorithms, whose performance is evaluated using Bayesian posterior diagnostics of energy response, spatial bias, and uncertainty calibration. The reconstruction fidelity is also analyzed across collision centralities and masked region sizes. This study establishes a general validation strategy for probabilistic reconstruction of missing detector information.

physics.data-an↗