arXiv Science⌕ Search

arXiv · 2609.38158

Goodness-of-fit for distributions on metric spaces

Abstract

We propose a general goodness-of-fit framework for distributions on separable metric spaces. Under suitable identifiability conditions, probability distributions are characterized by distance profiles, which motivates their use in goodness-of-fit testing, for simple and composite null hypotheses. For composite null hypotheses, parameter estimation is incorporated via a Bahadur-type expansion, and the asymptotic distribution of the empirical process for distance profiles is obtained under the null. We define test statistics based on this empirical process and derive their asymptotic null distributions. We further study the behavior of the proposed tests under fixed and local alternatives, establishing consistency results. Multiplier bootstrap procedures are developed, and their conditional asymptotic validity is established under both simple and composite null hypotheses. The methodology is illustrated with simulation studies and real-data applications for data on the sphere, hyperboloid, and simplex.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Diego Serrano, Eduardo García-Portugués, Ingrid Van Keilegom. 2026-09-29. Goodness-of-fit for distributions on metric spaces. https://arxiv.org/abs/2609.38158

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

On the Nonasymptotic Scaling Guarantee of Hyperparameter Estimation in Inhomogeneous, Weakly-Dependent Complex Network Dynamical Systems

Hierarchical Bayesian models are increasingly used in large, inhomogeneous complex network dynamical systems by modeling parameters as draws from a hyperparameter-governed distribution. However, theoretical guarantees for these estimates as the network population size grows have been lacking. A critical concern is that hyperparameter estimation may diverge for larger networks, undermining the model's reliability. Formulating the system's evolution in a measure transport perspective, we propose a theoretical framework for estimating hyperparameters with mean-type observations, which are prevalent in many scientific applications. Our primary contribution is a nonasymptotic bound for the deviation of estimate of hyperparameters in inhomogeneous complex network dynamical systems with respect to network population size, which is established for a general family of optimization algorithms within a fixed observation duration. For systems with independent nodes, we first establish a fixed-accuracy probabilistic hyperparameter bound and then use it to derive an explicit nonasymptotic hyperparameter bound. We subsequently extend both bounds to the more challenging and practically relevant setting of systems with weakly-dependent nodes. We validate our theoretical findings with numerical experiments on two representative models: a Susceptible-Infected-Susceptible model and a Spiking Neuronal Network model. In both cases, the results confirm that the estimation error decreases as the network population size increases, aligning with our theoretical guarantees. This research proposes the foundational theory to ensure that hierarchical Bayesian methods are statistically consistent for large-scale inhomogeneous systems, filling a gap in this area of theoretical research and justifying their application in practice.

math.ST↗

Intrinsic-dimension empirical Bernstein inequalities for bounded self-adjoint operators

Operator-valued concentration inequalities are foundational to the analysis of modern high-dimensional statistics and randomized algorithms. However, standard oracle bounds are frequently limited in practice: they require explicit a priori knowledge of the true variance, and often explicitly scale with the ambient dimension, rendering them vacuous for infinite-dimensional or heavily structured operators. Motivated by these challenges, we establish the first empirical Bennett and Bernstein inequalities for sums of independent, bounded, self-adjoint Hilbert-Schmidt operators. Our fully data-driven bounds replace the unknown variance with an empirical estimate and rely strictly on the intrinsic dimension rather than the ambient dimension. This structural shift yields computable, dimension-free guarantees with a sharper first-order asymptotic radius for non-isotropic random matrices and seamlessly extends to infinite-dimensional Hilbert spaces. We demonstrate that our empirical bounds achieve asymptotic sharpness with the best known oracle rates. Finally, as an independent byproduct, we derive novel empirical concentration guarantees for the intrinsic dimension itself.

math.ST↗

Bentkus-type asymptotic e-values

Asymptotic e-values are emerging as a powerful alternative to asymptotic p-values, particularly in post-hoc inference and multiple testing, where significance levels may be data-dependent. Existing asymptotic e-values, however, suffer from the ``missing factor,'' a scaling inefficiency resulting in overly conservative inference. Drawing on the framework of near-optimal concentration inequalities developed by Bentkus in the 2000s, we introduce Bentkus-type asymptotic e-values and prove that they successfully eliminate the missing factor. We also demonstrate both theoretically and empirically that Bentkus-type e-values consistently deliver sharper inference than existing alternatives, leading to tighter post-hoc confidence intervals and higher rejection rates in multiple testing procedures.

math.ST↗