arXiv ScienceSearch

arXiv subjects

Xiancheng Lin

Publications and source records attributed to Xiancheng Lin.

2 recordsLinked to original sources

UBSea: A Unified Community Detection Framework

Detecting communities in networks and graphs is an important task across many disciplines such as statistics, social science and engineering. There are generally three different kinds of mixing patterns for the case of two communities: assortative mixing, disassortative mixing and core-periphery structure. Modularity optimization is a classical way for fitting network models with communities. However, it can only deal with assortative mixing and disassortative mixing when the mixing pattern is known and fails to discover the core-periphery structure. In this paper, we extend modularity in a strategic way and propose a new framework based on Unified Bigroups Standadized Edge-count Analysis (UBSea). It can address all the formerly mentioned community mixing structures. In addition, this new framework is able to automatically choose the mixing type to fit the networks. Simulation studies show that the new framework has superb performance in a wide range of settings under the stochastic block model and the degree-corrected stochastic block model. We show that the new approach produces consistent estimate of the communities under a suitable signal-to-noise-ratio condition, for the case of a block model with two communities, for both undirected and directed networks. The new method is illustrated through applications to several real-world datasets.

stat.ME

High-Dimensional Clustering via Nearest-Neighbor Asymmetry

High-dimensional clustering often relies on geometric or local-similarity structure, but the dominant separation between groups may not always be location-based. Differences in dispersion can create asymmetric local-neighborhood patterns: points from a more dispersed component may be closer to points in a more concentrated component than to points from their own component. We turn this high-dimensional phenomenon into a clustering principle. The proposed method, NAC (Nearest-neighbor Asymmetry Clustering), constructs a directed $k$-nearest-neighbor graph and evaluates candidate partitions using two permutation-standardized statistics: a weighted within-edge statistic that captures overall within-cluster enrichment and a contrast statistic that captures asymmetric separation. The resulting objective combines these two standardized signals, allowing the method to adapt to different separation regimes without specifying a mixture model or a low-dimensional representation. We provide a population-level analysis showing how the two statistics target complementary nearest-neighbor patterns. Simulation studies across mean, scale, and combined location-scale differences show that NAC is competitive under location separation and especially effective when nearest-neighbor asymmetry is present; gene-expression applications further illustrate its usefulness in small-sample, high-dimensional clustering.

stat.ME