arXiv Science⌕ Search

arXiv · 2610.08711

Asymptotic Null Distributions of Moran's $I$ and Assortativity in Large Networks

Abstract

This study investigates the asymptotic behavior of two dependence measures defined on networks, Moran's $I$ statistic and Newman's assortativity, under the null hypothesis that a Gaussian node attribute $Y$ is independent of the network structure. We demonstrate that the structure of the network directly affects the convergence rate to normality of these measures as the size of the network increases. We further establish that, in some instances, the mean values of these dependence measures under the null hypothesis remain non-negligible asymptotically and must therefore be explicitly accounted for when calculating the test statistics. Applications to a variety of simulated and real networks also reveal that the normal approximation performs well only when the network is not strongly heterogeneous. Network topology determines both the convergence rate to normality and whether the limiting distribution is Gaussian. In dense networks whose degree heterogeneity does not vanish, we further show that assortativity can fail to be a valid test statistic even though Moran's $I$ remains well behaved, whereas a dominating node invalidates both.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Karin Ait Braham, Louis-Paul Rivest, Thierry Duchesne. 2026-10-06. Asymptotic Null Distributions of Moran's $I$ and Assortativity in Large Networks. https://arxiv.org/abs/2610.08711

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Geometric Identification and Consistent Estimation of Source Apportionment with Application to Greenland Summit Aerosols

Source apportionment, the attribution of observed multipollutant concentrations to underlying sources, can be cast as a non-negative matrix factorization (NMF) problem. Because NMF is non-unique, source apportionment imposes often unverifiable constraints such as structural sparsity or rotational tuning based on scientific knowledge, which may not be available in newer geographies or new sensor technology. Geometric NMF approaches offer a more data-driven route to identification, but many still rely on source profiles with arbitrary scalings, assume exact separability, and lack a framework for consistent estimation. We address these limitations by formally establishing statistical identifiability of the source apportionment matrix under a stochastic framework that replaces hard separability with soft probabilistic relaxations. We then present a scalable geometric algorithm to estimate this matrix and prove its consistency, to our knowledge the first such result requiring no exact sparsity, no parametric distributional assumptions, and accommodating spatio-temporal dependence. We apply this method in a setting where prior knowledge of sources is limited, so structural assumptions are challenging to impose. Over the last four decades, the Arctic has warmed four times faster than the rest of the planet, with anthropogenic aerosols among the major forcings. Greenland Summit Station has collected size-resolved particulate matter year-round since 2003, yet these records remain underutilized. Analyzing the 2003--2016 record with geometric NMF and HYSPLIT trajectories, we characterize long-range transport of anthropogenic and natural species into the Arctic.

math.ST↗

Shape without scale: an identifiability dichotomy for a bounded tail observed through a non-additive measurement kernel

A latent severity has a bounded lower tail with density of shape alpha and scale L. It is observed only through a fixed Markov kernel K that is biased and non-additive. The relative conditional spread of K diverges at the endpoint. Our sample is i.i.d. from the marginal Q alone, with no anchoring covariate or instrument. We prove a dichotomy. The shape index alpha is identifiable: for every admissible choice of the class constants, any two observationally equivalent members of a lean class share alpha, determined by a near-endpoint expansion of Q. The rate, namely L and the fixed-scale exceedance p_tau, does not survive. There exist admissible shared class constants and two members of a smaller regularity class whose observed laws coincide exactly. Across the pair alpha agrees, whereas L and p_tau move. A degenerate Le Cam two-point bound excludes any uniformly consistent estimator of either, and pointwise consistency fails at one member. Only the rate needs an anchor. We conjecture that a known kernel family with known edge map identifies the rate fiber by fiber if and only if the family satisfies a fixed-scale injectivity clause, and we prove the sufficiency direction. In surrogate safety, uncalibrated conflict data give the shape of near-crash risk, not its absolute rate.

math.ST↗

Tightness and Error Exponents of SDP with Logarithmically Many Communities

We study a semidefinite programming (SDP) relaxation for community recovery when the number of communities grows logarithmically. In the balanced stochastic block model with $n=km$ vertices, we consider the regime $k/\log m\toγ>0$, with edge probabilities $α\log m/m$ within communities and $β\log m/m$ across them, for fixed $α>β>0$. We derive the sharp asymptotic tightness boundary away from critical cases. When rare vertices cause tightness to fail while the bulk remains spectrally stable, the normalized matrix error of every near-optimal solution still vanishes. We prove matching high-probability exponents for this error and the normalized optimal objective gain, governed by the same local correction that determines tightness. Throughout this spectrally stable region, a single SDP solve followed by explicit rounding and refinement achieves exact community recovery above the information-theoretic threshold. This guarantee holds even when the planted community matrix is not an optimal solution to the SDP.

math.ST↗