arXiv Science⌕ Search

arXiv · 2610.09123

Sharp and Adaptive Cluster Recovery in Slowly Mixing Gaussian Hidden Markov Models

Abstract

We study cluster recovery in a two-state Gaussian hidden Markov model in the slowly mixing regime, where the transition probability $δ$ may vanish with the sample size. Temporal persistence creates long homogeneous segments and fundamentally changes the recovery problem relative to the i.i.d. Gaussian mixture benchmark. Rather than being purely adverse, label dependence provides exploitable structure: the statistical difficulty becomes concentrated near the rare transition points, giving rise to recovery regimes governed explicitly by their frequency $δ$. We first characterize, up to universal constants, the oracle Bayes risk for offline and online clustering. Both problems exhibit a nonstandard polynomial risk regime, but the sequential constraint induces an additional logarithmic localization cost, yielding a genuine online/offline gap. We also study fixed lag prediction with delayed label feedback and identify the effective memory horizon beyond which past labels no longer improve the optimal rate. When the signal direction is unknown and possibly high dimensional, the problem is governed by the effective signal strength \[ r_n^2 = \frac{\|θ\|^4/σ^4} {\|θ\|^2/σ^2+d/n}. \] We derive minimax lower bounds and construct fully adaptive polynomial time procedures attaining the sharp recovery rates. In particular, almost full recovery is possible precisely when $r_n^2\ggδ$. If $nδ\gtrsim\log n$, adaptation to both the signal $θ$ and the transition probability $δ$ incurs no first order loss, and the sharp exact recovery threshold (under the Hamming loss) is \[ r_n^2 = 2\log(nδ)\,(1+o(1)). \] Thus the effective complexity of the problem is governed by the number of latent transitions, rather than by the sample size itself.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ibrahim Kaddouri, Mohamed Ndaoud. 2026-10-06. Sharp and Adaptive Cluster Recovery in Slowly Mixing Gaussian Hidden Markov Models. https://arxiv.org/abs/2610.09123

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Geometric Identification and Consistent Estimation of Source Apportionment with Application to Greenland Summit Aerosols

Source apportionment, the attribution of observed multipollutant concentrations to underlying sources, can be cast as a non-negative matrix factorization (NMF) problem. Because NMF is non-unique, source apportionment imposes often unverifiable constraints such as structural sparsity or rotational tuning based on scientific knowledge, which may not be available in newer geographies or new sensor technology. Geometric NMF approaches offer a more data-driven route to identification, but many still rely on source profiles with arbitrary scalings, assume exact separability, and lack a framework for consistent estimation. We address these limitations by formally establishing statistical identifiability of the source apportionment matrix under a stochastic framework that replaces hard separability with soft probabilistic relaxations. We then present a scalable geometric algorithm to estimate this matrix and prove its consistency, to our knowledge the first such result requiring no exact sparsity, no parametric distributional assumptions, and accommodating spatio-temporal dependence. We apply this method in a setting where prior knowledge of sources is limited, so structural assumptions are challenging to impose. Over the last four decades, the Arctic has warmed four times faster than the rest of the planet, with anthropogenic aerosols among the major forcings. Greenland Summit Station has collected size-resolved particulate matter year-round since 2003, yet these records remain underutilized. Analyzing the 2003--2016 record with geometric NMF and HYSPLIT trajectories, we characterize long-range transport of anthropogenic and natural species into the Arctic.

math.ST↗

Shape without scale: an identifiability dichotomy for a bounded tail observed through a non-additive measurement kernel

A latent severity has a bounded lower tail with density of shape alpha and scale L. It is observed only through a fixed Markov kernel K that is biased and non-additive. The relative conditional spread of K diverges at the endpoint. Our sample is i.i.d. from the marginal Q alone, with no anchoring covariate or instrument. We prove a dichotomy. The shape index alpha is identifiable: for every admissible choice of the class constants, any two observationally equivalent members of a lean class share alpha, determined by a near-endpoint expansion of Q. The rate, namely L and the fixed-scale exceedance p_tau, does not survive. There exist admissible shared class constants and two members of a smaller regularity class whose observed laws coincide exactly. Across the pair alpha agrees, whereas L and p_tau move. A degenerate Le Cam two-point bound excludes any uniformly consistent estimator of either, and pointwise consistency fails at one member. Only the rate needs an anchor. We conjecture that a known kernel family with known edge map identifies the rate fiber by fiber if and only if the family satisfies a fixed-scale injectivity clause, and we prove the sufficiency direction. In surrogate safety, uncalibrated conflict data give the shape of near-crash risk, not its absolute rate.

math.ST↗

Spectral Analysis of Gaussian Integral Operators Arising in BHEP Tests

The Baringhaus-Henze-Epps-Pulley (BHEP) tests for multivariate normality are affine-invariant goodness-of-fit tests based on a Gaussian-weighted $L^2$ distance between empirical and Gaussian characteristic functions. In 1990, Henze and Zirkler expressed the limiting null distribution through the eigenvalues of an integral operator on the standard Gaussian space. In 1997, Henze and Wagner obtained a simpler covariance kernel and raised the problem of calculating the eigenvalues of the resulting operator on a Gaussian-weighted space. Although subsequent work treated the univariate case and numerical approximations in a few low dimensions, the complete all-dimensional spectral problem remained open. This paper determines both complete spectra for every dimension $d \in \mathbb{N}$ and every smoothing parameter $β> 0$. The two operators are shown to have the forms $\mathcal{X}_{β,d}^*\mathcal{X}_{β,d}$ and $\mathcal{X}_{β,d}\mathcal{X}_{β,d}^*$ for the same Hilbert-Schmidt operator $\mathcal{X}_{β,d}$. Consequently, their nonzero eigenvalues agree, including multiplicities, while the null space of the Henze-Zirkler operator is identified exactly. The Gaussian integral operator in the Henze-Wagner decomposition is diagonalized by Mehler's formula, and rotational symmetry confines the finite-rank correction to the sectors associated with spherical harmonics of degrees $0$, $1$, and $2$. The degree-$1$ and degree-$2$ eigenvalues are characterized by scalar transcendental equations, and the radial eigenvalues by an explicit pole-safe Fredholm determinant. The paper establishes nonnegativity, multiplicities, eigenfunction reconstruction, completeness, the trace identity, and a complete characterization of all exceptional pole cases.

math.ST↗