arXiv ScienceSearch

arXiv subjects

Huan Qing

Publications and source records attributed to Huan Qing.

At least 19 recordsLinked to original sources

Exact Community Recovery in Bipartite Networks

Community detection in bipartite networks is a fundamental problem in modern data analysis, with applications in recommendation systems, biological networks, and social network analysis. Unlike conventional unipartite graphs, bipartite networks consist of two distinct types of nodes with edges only connecting across types, so recovering latent communities requires estimating labels on the two node types. The stochastic co-blockmodel is a classical probabilistic framework for such networks, yet theoretical guarantees for exact community recovery in this setting remain limited, especially when the number of communities grows, the community sizes are unbalanced, or the degrees are heterogeneous. In this work, we prove that a simple spectral clustering algorithm based on the diagonal-deleted Gram matrix achieves exact recovery with high probability under mild conditions on sparsity, community balance, and the number of clusters. We further extend the result to the degree-corrected stochastic co-blockmodel, where each node carries its own degree heterogeneity parameter, and show that a row-normalized version of the same algorithm maintains the exact recovery guarantee. Extensive experiments validate our theoretical findings.

cs.SI

A Subsampled Davis-Kahan Bound for Large-Scale Eigenspace Estimation

The Davis-Kahan theorem is a fundamental tool in spectral analysis, providing quantitative control over the distance between the eigenspaces of a symmetric matrix and its perturbation. However, when the matrix dimension is large, computing leading eigenvectors is computationally expensive, limiting the practical use of spectral methods in modern large-scale applications. This paper addresses this problem by proposing an independent Bernoulli sampling scheme and proves that the leading left singular vectors of the subsampled matrix faithfully approximate the target subspace of a low-rank symmetric matrix. Our main result is a subsampled Davis-Kahan bound that gives an explicit error bound depending directly on the sampling probability. The bound reveals the trade-off: the computational cost scales linearly with the sampling probability, while the statistical error scales as the inverse square root of the sampling probability. Our result thus extends the Davis-Kahan theorem to the subsampled setting, enabling scalable spectral analysis of large-scale symmetric matrices.

stat.ML

Latent class analysis by regularized spectral clustering

The latent class model is a highly effective tool in the analysis of categorical data from social, psychological, and behavioral sciences, where populations often share hidden common characteristics. In this article, we introduce two new algorithms for estimating the parameters of a latent class model for ordered categorical data with polytomous responses. These algorithms are based on a newly defined regularized Laplacian matrix derived from the response matrix. We provide theoretical convergence rates for our algorithms by considering a sparsity parameter and demonstrate that under a mild condition on data's sparsity, our algorithms yield consistent latent class analysis. Furthermore, we introduce a metric to assess the strength of latent class analysis and develop procedures based on this metric to determine the optimal number of latent classes for real-world ordered categorical data. Extensive simulation experiments demonstrate the efficiency and accuracy of our algorithms, and we demonstrate their practical application to real-world ordered categorical data with promising results.

cs.LG

Directed mixed membership stochastic blockmodel

Mixed membership modeling for undirected networks has been extensively explored in network science over the past few years. Despite the substantial progress made for undirected cases, handling mixed membership structures in directed networks continues to pose substantial difficulties. To address this gap, we introduce the Directed Mixed Membership Stochastic Blockmodel (DiMMSB), a novel framework tailored for directed networks with overlapping communities. A key feature of DiMMSB is its ability to treat the row and column nodes of the adjacency matrix as distinct entities, each potentially following its own community organization. Building on this model, we develop DiSP, an efficient spectral procedure to estimate mixed memberships for both sets of nodes. Through delicate analysis, we derive node-specific error bounds of DiSP under mild sparsity conditions. Simulation results support the theoretical results, demonstrating that DiSP achieves lower error rates and faster computation than its competitor. Moreover, applications to real data highlight DiSP's effectiveness in uncovering asymmetric structural patterns.

stat.ML

Mixed Membership sub-Gaussian Models

The Gaussian mixture model is widely used in unsupervised learning, owing to its simplicity and interpretability. However, a fundamental limitation of the classical Gaussian mixture model is that it forces each observation to belong to exactly one component. In many practical applications, such as genetics, social network analysis, and text mining, an observation may naturally belong to multiple components or exhibit partial membership in several latent components. To overcome this limitation, we propose the mixed membership sub-Gaussian model, which extends the classical Gaussian mixture framework by allowing each observation to belong to multiple components. This model inherits the interpretability of the classical Gaussian mixture model while offering greater flexibility for capturing complex overlapping structures. We develop an efficient spectral algorithm to estimate the mixed membership of each individual observation, and under mild separation conditions on the component centres, we prove that the estimation error of the per-individual membership vector can be made arbitrarily small with high probability. To our knowledge, this is the first work to provide a computationally efficient estimator with such a vanishing-error guarantee for a mixed-membership extension of the Gaussian mixture model. Extensive experimental studies demonstrate that our method outperforms existing approaches that ignore mixed memberships.

stat.ML

Fast estimation of Gaussian mixture components via centering and singular value thresholding

Estimating the number of components is a fundamental challenge in unsupervised learning, particularly when dealing with high-dimensional data with many components or severely imbalanced component sizes. This paper addresses this challenge for classical Gaussian mixture models. The proposed estimator is simple: center the data, compute the singular values of the centered matrix, and count those above a threshold. No iterative fitting, no likelihood calculation, and no prior knowledge of the number of components are required. We prove that, under a mild separation condition on the component centers, the estimator consistently recovers the true number of components. The result holds in high-dimensional settings where the dimension can be much larger than the sample size. It also holds when the number of components grows to the smaller of the dimension and the sample size, even under severe imbalance among component sizes. Computationally, the method is extremely fast: for example, it processes ten million samples in one hundred dimensions within one minute. Extensive experimental studies confirm its accuracy in challenging settings such as high dimensionality, many components, and severe class imbalance.

stat.ML

Finding Core Balanced Modules in Statistically Validated Stock Networks

Traditional threshold-based stock networks suffer from subjective parameter selection and inherent limitations: they constrain relationships to binary representations, failing to capture both correlation strength and negative dependencies. To address this, we introduce statistically validated correlation networks that retain only statistically significant correlations via a rigorous t-test of Pearson coefficients. We then propose a novel structure termed the largest strong correlation balanced module (LSCBM), defined as the maximum-size group of stocks with structural balance (i.e., positive edge-sign products for all triplets) and strong pairwise correlations. This balance condition ensures stable relationships, thus facilitating potential hedging opportunities through negative edges. Theoretically, within a random signed graph model, we establish LSCBM's asymptotic existence, size scaling, and multiplicity under various parameter regimes. To detect LSCBM efficiently, we develop MaxBalanceCore, a heuristic algorithm that leverages network sparsity. Simulations validate its efficiency, demonstrating scalability to networks of up to 10,000 nodes within tens of seconds. Empirical analysis demonstrates that LSCBM identifies core market subsystems that dynamically reorganize in response to economic shifts and crises. In the Chinese stock market (2013-2024), LSCBM's size surges during high-stress periods (e.g., the 2015 crash) and contracts during stable or fragmented regimes, while its composition rotates annually across dominant sectors (e.g., Industrials and Financials).

econ.GN

Individual-heterogeneous sub-Gaussian Mixture Models

The classical Gaussian mixture model assumes homogeneity within clusters, an assumption that often fails in real-world data where observations naturally exhibit varying scales or intensities. To address this, we introduce the individual-heterogeneous sub-Gaussian mixture model, a flexible framework that assigns each observation its own heterogeneity parameter, thereby explicitly capturing the heterogeneity inherent in practical applications. Built upon this model, we propose an efficient spectral method that provably achieves exact recovery of the true cluster labels under mild separation conditions, even in high-dimensional settings where the number of features far exceeds the number of samples. Numerical experiments on both synthetic and real data demonstrate that our method consistently outperforms existing clustering algorithms, including those designed for classical Gaussian mixture models.

stat.ML

Grade of membership analysis for multi-layer ordinal categorical data

Consider a group of individuals (subjects) participating in the same psychological tests with numerous questions (items) at different times, where the choices of each item have an implicit ordering. The observed responses can be recorded in multiple response matrices over time, named multi-layer ordinal categorical data, where layers refer to time points. Assuming that each subject has a common mixed membership shared across all layers, enabling it to be affiliated with multiple latent classes with varying weights, the objective of the grade of membership (GoM) analysis is to estimate these mixed memberships from the data. When the test is conducted only once, the data becomes traditional single-layer ordinal categorical data. The GoM model is a popular choice for describing single-layer categorical data with a latent mixed membership structure. However, GoM cannot handle multi-layer ordinal categorical data. In this work, we propose a new model, multi-layer GoM, which extends GoM to multi-layer ordinal categorical data. To estimate the common mixed memberships, we propose a new approach, GoM-DSoG, based on a debiased sum of Gram matrices. We establish GoM-DSoG's per-subject convergence rate under the multi-layer GoM model. Our theoretical results suggest that fewer no-responses, more subjects, more items, and more layers are beneficial for GoM analysis. We also propose an approach to select the number of latent classes. Extensive experimental studies verify the theoretical findings and show GoM-DSoG's superiority over its competitors, as well as the accuracy of our method in determining the number of latent classes.

stat.ME

Goodness-of-fit test for multi-layer stochastic block models

Community detection in multi-layer networks is a fundamental task in complex network analysis across various areas like social, biological, and computer sciences. However, most existing algorithms assume that the number of communities is known in advance, which is usually impractical for real-world multi-layer networks. To address this limitation, we develop a novel goodness-of-fit test for the popular multi-layer stochastic block model based on a normalized aggregation of layer-wise adjacency matrices. Under the null hypothesis that a candidate community count is correct, we establish the asymptotic normality of the test statistic using recent advances in random matrix theory; conversely, we prove its divergence when the model is underfitted. This dual theoretical foundations enable two computationally efficient sequential testing algorithms to consistently determine the number of communities without prior knowledge. Numerical experiments on simulated and real-world multi-layer networks demonstrate the accuracy and efficiency of our approaches in estimating the number of communities.

stat.ME

How many asymmetric communities are there in multi-layer directed networks?

Estimating the asymmetric numbers of communities in multi-layer directed networks is a challenging problem due to the multi-layer structures and inherent directional asymmetry, leading to possibly different numbers of sender and receiver communities. This work addresses this issue under the multi-layer stochastic co-block model, a model for multi-layer directed networks with distinct community structures in sending and receiving sides, by proposing a novel goodness-of-fit test. The test statistic relies on the deviation of the largest singular value of an aggregated normalized residual matrix from the constant 2. The test statistic exhibits a sharp dichotomy: Under the null hypothesis of correct model specification, its upper bound converges to zero with high probability; under underfitting, the test statistic itself diverges to infinity. With this property, we develop a sequential testing procedure that searches through candidate pairs of sender and receiver community numbers in a lexicographic order. The process stops at the smallest such pair where the test statistic drops below a decaying threshold. For robustness, we also propose a ratio-based variant algorithm, which detects sharp changes in the sequence of test statistics by comparing consecutive candidates. Both methods are proven to consistently determine the true numbers of sender and receiver communities under the multi-layer stochastic co-block model.

math.ST

Goodness-of-Fit Tests for Latent Class Models with Ordinal Categorical Data

Ordinal categorical data are widely collected in psychology, education, and other social sciences, appearing commonly in questionnaires, assessments, and surveys. Latent class models provide a flexible framework for uncovering unobserved heterogeneity by grouping individuals into homogeneous classes based on their response patterns. A fundamental challenge in applying these models is determining the number of latent classes, which is unknown and must be inferred from data. In this paper, we propose one test statistic for this problem. The test statistic centers the largest singular value of a normalized residual matrix by a simple sample-size adjustment. Under the null hypothesis that the candidate number of latent classes is correct, its upper bound converges to zero in probability. Under an under-fitted alternative, the statistic itself exceeds a fixed positive constant with probability approaching one. This sharp dichotomous behavior of the test statistic yields two sequential testing algorithms that consistently estimate the true number of latent classes. Extensive experimental studies confirm the theoretical findings and demonstrate their accuracy and reliability in determining the number of latent classes.

stat.ML

Overlapping community detection in weighted networks

Over the past decade, community detection in overlapping un-weighted networks, where nodes can belong to multiple communities, has been one of the most popular topics in modern network science. However, community detection in overlapping weighted networks, where edge weights can be any real value, remains challenging. In this article, we propose a generative model called the weighted degree-corrected mixed membership (WDCMM) model to model such weighted networks. This model adopts the same factorization for the expectation of the adjacency matrix as the previous degree-corrected mixed membership (DCMM) model. Our WDCMM extends the DCMM from un-weighted networks to weighted networks by allowing the elements of the adjacency matrix to be generated from distributions beyond Bernoulli. We first address the community membership estimation of the model by applying a spectral algorithm and establishing a theoretical guarantee of consistency. Then, we propose overlapping weighted modularity to measure the quality of overlapping community detection for both assortative and dis-assortative weighted networks. To determine the number of communities, we incorporate the algorithm into the proposed modularity. We demonstrate the advantages of the model and the modularity through applications to simulated data and real-world networks.

cs.SI

Mixed membership estimation for categorical data with weighted responses

The Grade of Membership (GoM) model, which allows subjects to belong to multiple latent classes, is a powerful tool for inferring latent classes in categorical data. However, its application is limited to categorical data with nonnegative integer responses, as it assumes that the response matrix is generated from Bernoulli or Binomial distributions, making it inappropriate for datasets with continuous or negative weighted responses. To address this, this paper proposes a novel model named the Weighted Grade of Membership (WGoM) model. Our WGoM is more general than GoM because it relaxes GoM's distribution constraint by allowing the response matrix to be generated from distributions like Bernoulli, Binomial, Normal, and Uniform as long as the expected response matrix has a block structure related to subjects' mixed memberships under the distribution. We show that WGoM can describe any response matrix with finite distinct elements. We then propose an algorithm to estimate the latent mixed memberships and other WGoM parameters. We derive the error bounds of the estimated parameters and show that the algorithm is statistically consistent. We also propose an efficient method for determining the number of latent classes $K$ for categorical data with weighted responses by maximizing fuzzy weighted modularity. The performance of our methods is validated through both synthetic and real-world datasets. The results demonstrate the accuracy and efficiency of our algorithm for estimating latent mixed memberships, as well as the high accuracy of our method for estimating $K$, indicating their high potential for practical applications.

cs.SI

Joint estimation of asymmetric community numbers in directed networks

Community detection in directed networks is a central task in network analysis. Unlike undirected networks, directed networks encode inherently asymmetric relationships, giving rise to sender and receiver roles that may each follow distinct community organizations with possibly different numbers of communities. Estimating these two community counts simultaneously is therefore considerably more challenging than in the undirected setting, yet it is essential for faithful model specification and reliable downstream inference. This work addresses this challenge within the stochastic co-block model (ScBM), a powerful statistical framework for capturing asymmetric relational structures inherent in directed networks. We propose a novel goodness-of-fit test based on the deviation of the largest singular value of a normalized residual matrix from the constant value 2. We show that the upper bound of this test statistic converges to zero under the null hypothesis, while this statistic goes to infinity if the true model has finer communities than hypothesized. Leveraging this tail bounds behavior, we develop an efficient sequential testing algorithm that lexicographically explores candidate community number pairs. To enhance robustness in practical settings, we further introduce a ratio-based variant that detects the transition point in the test statistic sequence. We rigorously show both algorithms' consistency in recovering the true sender and receiver community counts under ScBM. Numerical experiments demonstrate the accuracy and robustness of our methods in estimating community numbers across diverse ScBM settings. %To our knowledge, this work presents the first theoretically guaranteed approach for jointly estimating the numbers of sender and receiver communities within the ScBM framework, providing a critical tool for reliable directed network analysis.

stat.ME

Community detection in multi-layer networks by regularized debiased spectral clustering

Community detection is a crucial problem in the analysis of multi-layer networks. While regularized spectral clustering methods using the classical regularized Laplacian matrix have shown great potential in handling sparse single-layer networks, to our knowledge, their potential in multi-layer network community detection remains unexplored. To address this gap, in this work, we introduce a new method, called regularized debiased sum of squared adjacency matrices (RDSoS), to detect communities in multi-layer networks. RDSoS is developed based on a novel regularized Laplacian matrix that regularizes the debiased sum of squared adjacency matrices. In contrast, the classical regularized Laplacian matrix typically regularizes the adjacency matrix of a single-layer network. Therefore, at a high level, our regularized Laplacian matrix extends the classical one to multi layer networks. We establish the consistency property of RDSoS under the multi-layer stochastic block model (MLSBM) and further extend RDSoS and its theoretical results to the degree-corrected version of the MLSBM model. Additionally, we introduce a sum of squared adjacency matrices modularity (SoS-modularity) to measure the quality of community partitions in multi-layer networks and estimate the number of communities by maximizing this metric. Our methods offer promising applications for predicting gene functions, improving recommender systems, detecting medical insurance fraud, and facilitating link prediction. Experimental results demonstrate that our methods exhibit insensitivity to the selection of the regularizer, generally outperform state-of-the-art techniques, uncover the assortative property of real networks, and that our SoS-modularity provides a more accurate assessment of community quality compared to the average of the Newman-Girvan modularity across layers.

stat.ME

Discovering overlapping communities in multi-layer directed networks

Community detection in multi-layer undirected networks has attracted considerable attention in recent years. However, multi-layer directed networks are common in the real world, and existing community detection methods often either ignore the asymmetric structure in multi-layer directed networks or assume that every node solely belongs to a single community, significantly limiting their applicability to overlapping multi-layer directed networks, where nodes can belong to multiple communities simultaneously. To fill this gap, this article explores the challenging problem of detecting overlapping communities in multi-layer directed networks. Our goal is to understand the underlying asymmetric overlapping community structure by analyzing the mixed memberships of nodes. We introduce a novel multi-layer mixed membership stochastic co-block model (multi-layer MM-ScBM) to model overlapping multi-layer directed networks. We develop a spectral procedure to estimate nodes' memberships in both sending and receiving patterns. Our method uses a successive projection algorithm on a few leading eigenvectors of two debiased aggregation matrices. To our knowledge, this is the first work to detect asymmetric overlapping communities in multi-layer directed networks. We demonstrate the consistent estimation properties of our method by providing per-node error rates under the multi-layer MM-ScBM framework. Our theoretical analysis reveals that increasing the overall sparsity, the number of nodes, or the number of layers can improve the accuracy of overlapping community detection. Extensive numerical experiments validate these theoretical findings. We also apply our method to one real-world multi-layer directed network, gaining insightful results.

cs.SI

Community detection by spectral methods in multi-layer networks

Community detection in multi-layer networks is a crucial problem in network analysis. In this paper, we analyze the performance of two spectral clustering algorithms for community detection within the framework of the multi-layer degree-corrected stochastic block model (MLDCSBM) framework. One algorithm is based on the sum of adjacency matrices, while the other utilizes the debiased sum of squared adjacency matrices. We also provide their accelerated versions through subsampling to handle large-scale multi-layer networks. We establish consistency results for community detection of the two proposed methods under MLDCSBM as the size of the network and/or the number of layers increases. Our theorems demonstrate the advantages of utilizing multiple layers for community detection. Our analysis also indicates that spectral clustering with the debiased sum of squared adjacency matrices is generally superior to spectral clustering with the sum of adjacency matrices. Furthermore, we provide a strategy to estimate the number of communities in multi-layer networks by maximizing the averaged modularity. Substantial numerical simulations demonstrate the superiority of our algorithm employing the debiased sum of squared adjacency matrices over existing methods for community detection in multi-layer networks, the high computational efficiency of our accelerated algorithms for large-scale multi-layer networks, and the high accuracy of our strategy in estimating the number of communities. Finally, the analysis of several real-world multi-layer networks yields meaningful insights.

cs.SI