arXiv Science⌕ Search

arXiv · 2610.06873

Entropy-Based Goodness-of-Fit Tests for Generalized $κ$ -Gaussian Distributions

Abstract

We introduce this class of multivariate Kappa-Gaussian distributions by maximizing a type of Kaniadakis $κ$ entropy subject to certain first- and second-moment constraints. This differs slightly from the approach of Rényi and Tsallis, because the $κ$-entropy arises from a relativity-consistent deformation of the exponential and logarithmic functions and yields a family of distributions whose tails follow a power law, a family that actually reverts to the standard Gaussian distribution as $κ\to 0$. We then create a $k$-nearest-neighbor (NN) estimator for $κ$-entropy and describe it using the difference between two NN functionals involving sums of powers. Next, we establish $L^2$-consistency and asymptotic normality, using the well-known subadditive Euclidean functional and moment expansion framework for NN entropy estimators. Furthermore, using a scaling identity for the $κ$-exponential, we derive a closed-form formula for the entropy of the $κ$-Gaussian distribution itself. From this, we determine the exact threshold value $κ< 2/(m+2)$ at which the distribution preserves finite variance; this is analogous to the ``degrees of freedom $>2$'' condition in the Student's-$t$ distribution. With these components, we construct a non-parametric goodness-of-fit test and then verify it through simulations. Finally, we mention the Sharma--Mittal entropy, which combines the Rényi and Tsallis entropies as a natural direction for future work.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mehmet Siddik Cadirci, Konstantinos Zografos. 2026-09-08. Entropy-Based Goodness-of-Fit Tests for Generalized $κ$ -Gaussian Distributions. https://arxiv.org/abs/2610.06873

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Geometric Identification and Consistent Estimation of Source Apportionment with Application to Greenland Summit Aerosols

Source apportionment, the attribution of observed multipollutant concentrations to underlying sources, can be cast as a non-negative matrix factorization (NMF) problem. Because NMF is non-unique, source apportionment imposes often unverifiable constraints such as structural sparsity or rotational tuning based on scientific knowledge, which may not be available in newer geographies or new sensor technology. Geometric NMF approaches offer a more data-driven route to identification, but many still rely on source profiles with arbitrary scalings, assume exact separability, and lack a framework for consistent estimation. We address these limitations by formally establishing statistical identifiability of the source apportionment matrix under a stochastic framework that replaces hard separability with soft probabilistic relaxations. We then present a scalable geometric algorithm to estimate this matrix and prove its consistency, to our knowledge the first such result requiring no exact sparsity, no parametric distributional assumptions, and accommodating spatio-temporal dependence. We apply this method in a setting where prior knowledge of sources is limited, so structural assumptions are challenging to impose. Over the last four decades, the Arctic has warmed four times faster than the rest of the planet, with anthropogenic aerosols among the major forcings. Greenland Summit Station has collected size-resolved particulate matter year-round since 2003, yet these records remain underutilized. Analyzing the 2003--2016 record with geometric NMF and HYSPLIT trajectories, we characterize long-range transport of anthropogenic and natural species into the Arctic.

math.ST↗

Shape without scale: an identifiability dichotomy for a bounded tail observed through a non-additive measurement kernel

A latent severity has a bounded lower tail with density of shape alpha and scale L. It is observed only through a fixed Markov kernel K that is biased and non-additive. The relative conditional spread of K diverges at the endpoint. Our sample is i.i.d. from the marginal Q alone, with no anchoring covariate or instrument. We prove a dichotomy. The shape index alpha is identifiable: for every admissible choice of the class constants, any two observationally equivalent members of a lean class share alpha, determined by a near-endpoint expansion of Q. The rate, namely L and the fixed-scale exceedance p_tau, does not survive. There exist admissible shared class constants and two members of a smaller regularity class whose observed laws coincide exactly. Across the pair alpha agrees, whereas L and p_tau move. A degenerate Le Cam two-point bound excludes any uniformly consistent estimator of either, and pointwise consistency fails at one member. Only the rate needs an anchor. We conjecture that a known kernel family with known edge map identifies the rate fiber by fiber if and only if the family satisfies a fixed-scale injectivity clause, and we prove the sufficiency direction. In surrogate safety, uncalibrated conflict data give the shape of near-crash risk, not its absolute rate.

math.ST↗

Tightness and Error Exponents of SDP with Logarithmically Many Communities

We study a semidefinite programming (SDP) relaxation for community recovery when the number of communities grows logarithmically. In the balanced stochastic block model with $n=km$ vertices, we consider the regime $k/\log m\toγ>0$, with edge probabilities $α\log m/m$ within communities and $β\log m/m$ across them, for fixed $α>β>0$. We derive the sharp asymptotic tightness boundary away from critical cases. When rare vertices cause tightness to fail while the bulk remains spectrally stable, the normalized matrix error of every near-optimal solution still vanishes. We prove matching high-probability exponents for this error and the normalized optimal objective gain, governed by the same local correction that determines tightness. Throughout this spectrally stable region, a single SDP solve followed by explicit rounding and refinement achieves exact community recovery above the information-theoretic threshold. This guarantee holds even when the planted community matrix is not an optimal solution to the SDP.

math.ST↗