arXiv ScienceSearch

arXiv · 2210.12580

A dichotomous behavior of Guttman-Kaiser criterion from equi-correlated normal population

Abstract

We consider a $p$-dimensional, centered normal population such that all variables have a positive variance $σ^2$ and any correlation coefficient between different variables is a given nonnegative constant $ρ<1$. Suppose that both the sample size $n$ and population dimension $p$ tend to infinity with $p/n \to c>0$. We prove that the limiting spectral distribution of a sample correlation matrix is Marčenko-Pastur distribution of index $c$ and scale parameter $1-ρ$. By the limiting spectral distributions, we rigorously show the limiting behavior of widespread stopping rules Guttman-Kaiser criterion and cumulative-percentage-of-variation rule in PCA and EFA. As a result, we establish the following dichotomous behavior of Guttman-Kaiser criterion when both $n$ and $p$ are large, but $p/n$ is small: (1) the criterion retains a small number of variables for $ρ>0$, as suggested by Kaiser, Humphreys, and Tucker [Kaiser, H. F. (1992). On Cliff's formula, the Kaiser-Guttman rule and the number of factors. Percept. Mot. Ski. 74]; and (2) the criterion retains $p/2$ variables for $ρ=0$, as in a simulation study [Yeomans, K. A. and Golder, P. A. (1982). The Guttman-Kaiser criterion as a predictor of the number of common factors. J. Royal Stat. Soc. Series D. 31(3)].

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yohji Akama, Atina Husnaqilati. 2023-01-09. A dichotomous behavior of Guttman-Kaiser criterion from equi-correlated normal population. https://arxiv.org/abs/2210.12580

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PDE-constrained inverse problems at the $\sqrt{n}$ rate via debiased physics-informed neural networks

We study the problem of estimating unknown parameters in PDE-constrained inverse problems from noisy observations, where the PDE solution is approximated using Physics-Informed Neural Networks (PINNs). While PINNs have demonstrated remarkable empirical success, existing estimators often inherit the slow nonparametric convergence rate of the neural-network solution, leading to biased and statistically inefficient inference for the finite-dimensional parameters of interest. To address this, we propose a two-step debiased estimation procedure that combines neural-network-based nonparametric estimation with an influence-function-based bias correction. By eliminating the first-order sensitivity of the estimator to errors in the nuisance function, our procedure yields a $\sqrt{n}$-consistent and asymptotically normal estimator without requiring undersmoothing of the neural network component. We further extend this framework to Bayesian inference by replacing the original likelihood with a debiased quasi-likelihood and establish a Bernstein-von Mises theorem showing that the resulting posterior contracts at the $\sqrt{n}$-rate with an asymptotic covariance matching that of the frequentist estimator. As a by-product of our analysis, we establish near-minimax optimal convergence rates for estimating a nonparametric regression function and its derivatives in Sobolev spaces using neural networks. Extensive numerical experiments corroborate our theoretical findings and demonstrate the necessity of the proposed debiasing procedure for valid statistical inference in PDE-constrained inverse problems.

math.ST

On spectral gap decomposition for Markov chains

Multiple works regarding convergence analysis of Markov chains have led to spectral gap decomposition formulas of the form \[ \mathrm{Gap}(S) \geq c_0 \left[\inf_z \mathrm{Gap}(Q_z)\right] \mathrm{Gap}(\bar{S}), \] where $c_0$ is a constant, $\mathrm{Gap}$ denotes the right spectral gap of a reversible Markov operator, $S$ is the Markov transition kernel (Mtk) of interest, $\bar{S}$ is an idealized or simplified version of $S$, and $\{Q_z\}$ is a collection of Mtks characterizing the differences between $S$ and $\bar{S}$. This type of relationship has been established in various contexts, including: 1. decomposition of Markov chains based on a finite cover of the state space, 2. hybrid Gibbs samplers, and 3. spectral independence and localization schemes. We show that multiple key decomposition results across these domains can be connected within a unified framework, rooted in a simple sandwich structure of $S$. Within the general framework, we establish new instances of spectral gap decomposition for hybrid hit-and-run samplers and hybrid data augmentation algorithms with two intractable conditional distributions. Additionally, we explore several other properties of the sandwich structure, and derive extensions of the spectral gap decomposition formula.

math.ST

Stable Central Limit Theorems for Discrete-Time Lag Martingale Difference Arrays: Applications to Dynamic Causal Inference

Recent work in dynamic causal inference introduced a class of discrete-time stochastic processes that generalize martingale difference sequences and arrays as follows: the random variates in each sequence have expectation zero given certain lagged filtrations but not given the natural filtration. We formalize this class of stochastic processes and prove stable central limit theorems (CLTs) via martingale-coboundary decomposition, leveraging the classical martingale CLT. We develop a variety of sufficient conditions, including conditions under which the limiting variance has a simple form that depends on variances and covariances of neighboring variates. We demonstrate the application of these results to inference for time-averaged treatment effects in switchback designs and present a simulation study supporting their validity. The CLTs enable various extensions to existing methodology for design-based approaches to dynamic causal inference, including time-lagged effects, random limiting variances, cross-unit dependence, and vector-valued estimands.

math.ST