arXiv ScienceSearch

arXiv subjects

Dan Zhuang

Publications and source records attributed to Dan Zhuang.

6 recordsLinked to original sources

Spectral Dynamics of DeepWalk Embeddings for Dynamic Network Change-Point Detection

Dynamic networks describe evolving relational systems in which abrupt structural changes may signal anomalous events or important transitions. Detecting such changes requires distinguishing genuine structural signals from fluctuations in network observations and learned representations. We propose a DeepWalk-based framework for detecting and localizing structural changes in dynamic networks. For each snapshot, we learn low-dimensional node embeddings, align them to a fixed reference using orthogonal Procrustes transformations, and aggregate them by mean pooling into comparable graph-level vectors. We then construct a multivariate cumulative sum (CUSUM) scan statistic to identify changes in the embedding mean. Under stated regularity conditions, we derive bounds on stochastic fluctuations under the null hypothesis and sufficient conditions for reliable detection and consistent localization under the alternative. The analysis connects detection performance with network sparsity, spectral separation, embedding dimension, and change magnitude, clarifying the balance between structural signals and representation variability. Simulation studies demonstrate the effectiveness of the proposed method in detecting and localizing structural changes across a range of dynamic network settings. An application to an international trade network further illustrates its practical usefulness in identifying shifts in trade allocation.

stat.ME

Change-Point Detection for Heterogeneous High-Dimensional Functional Time Series

High-dimensional functional panels consist of temporally ordered curves observed across many subjects and naturally exhibit heterogeneous structural changes. Under sparse subject-level break signals or opposite-signed shifts, traditional mean-aggregated CUSUM procedures may suffer noticeable power loss due to signal attenuation or cancellation induced by cross-sectional averaging. We propose a novel Energy--PE statistic, which combines subject-wise squared CUSUM energy aggregation with a generalized power-enhancement component. The energy aggregation preserves subject-level evidence under sign-heterogeneous changes, while the power-enhancement component improves sensitivity to sparse weak break signals. Under regularity conditions, we establish the asymptotic behavior of the proposed statistic. We further incorporate a latent group structure and an information-criterion-based clustering algorithm to estimate the unknown group number and membership for heterogeneous break points. Numerical studies and an intraday stock application demonstrate that Energy--PE controls size, improves power under sparse and sign-heterogeneous alternatives, and yields interpretable post-test summaries.

stat.ME

Semiparametric Elliptical Mixture Clustering for High-Dimensional Data

Clustering high-dimensional data is especially challenging when cluster distributions are heavy tailed and only approximately elliptical. Existing high-dimensional methods are largely built for Gaussian or other light-tailed models, whereas classical robust elliptical procedures are mostly low dimensional or rely on fully parametric radial families. We propose a semiparametric elliptical mixture clustering framework with cluster-specific centers, an unknown common radial generator, and a common sparse precision-shape matrix, together with a data-driven rule for selecting the number of clusters. A generalized expectation-maximization (GEM) algorithm is developed by combining transformed-radius estimation of the radial generator, radial-score center updates, and a Tyler-POET-GLASSO update for the common precision-shape matrix. The method avoids specifying a parametric radial family and remains computationally feasible in high dimensions. We establish high-dimensional consistency for the estimated model components and the excess misclustering error. Simulation studies and a handwritten-digit application demonstrate the competitive performance and robustness of the proposed method, particularly in heavy-tailed elliptical settings.

stat.ME

Sparse $K$-spatial-median clustering for high-dimensional data

We propose a robust clustering framework for high-dimensional data with heavy tails and a large fraction of irrelevant variables. The method replaces the mean updates of Lloyd's $K$-means with \emph{spatial medians} to enhance robustness. For the assignment step, it admits either a Euclidean rule for computational simplicity or a robust Mahalanobis-type metric constructed from the spatial sign covariance matrix to account for heterogeneous scales and feature dependence. To handle the $p \gg n$ regime, we further introduce a simple \emph{hard feature-exclusion} mechanism that removes weakly separating dimensions based on across-center dispersion, with the exclusion threshold selected automatically via a permutation-based Gap criterion. Simulation studies under correlated Gaussian and multivariate $t$ models demonstrate that the proposed approach provides competitive clustering accuracy and improved stability relative to $K$-means and sparse $K$-means baselines.

stat.ME

Adaptive Test for High Dimensional Quantile Regression

Testing high-dimensional quantile regression coefficients is crucial, as tail quantiles often reveal more than the mean in many practical applications. Nevertheless, the sparsity pattern of the alternative hypothesis is typically unknown in practice, posing a major challenge. To address this, we propose an adaptive test that remains powerful across both sparse and dense alternatives.We first establish the asymptotic independence between the max-type test statistic proposed by \citet{tang2022conditional} and the sum-type test statistic introduced by \citet{chen2024hypothesis}. Building on this result, we propose a Cauchy combination test that effectively integrates the strengths of both statistics and achieves robust performance across a wide range of sparsity levels. Simulation studies and real data applications demonstrate that our proposed procedure outperforms existing methods in terms of both size control and power.

stat.ME

Spatial Sign based Direct Sparse Linear Discriminant Analysis for High Dimensional Data

This paper investigates the robust linear discriminant analysis (LDA) problem with elliptical distributions in high-dimensional data. We propose a robust classification method, named SSLDA, that is intended to withstand heavy-tailed distributions. We demonstrate that SSLDA achieves an optimal convergence rate in terms of both misclassification rate and estimate error. Our theoretical results are further confirmed by extensive numerical experiments on both simulated and real datasets. Compared with current approaches, the SSLDA method offers superior improved finite sample performance and notable robustness against heavy-tailed distributions.

stat.ME