arXiv ScienceSearch

arXiv · 1710.05291

Testing for Principal Component Directions under Weak Identifiability

Abstract

We consider the problem of testing, on the basis of a $p$-variate Gaussian random sample, the null hypothesis ${\cal H}_0: {\pmb θ}_1= {\pmb θ}_1^0$ against the alternative ${\cal H}_1: {\pmb θ}_1 \neq {\pmb θ}_1^0$, where ${\pmb θ}_1$ is the "first" eigenvector of the underlying covariance matrix and ${\pmb θ}_1^0$ is a fixed unit $p$-vector. In the classical setup where eigenvalues $λ_1>λ_2\geq \ldots\geq λ_p$ are fixed, the Anderson (1963) likelihood ratio test (LRT) and the Hallin, Paindaveine and Verdebout (2010) Le Cam optimal test for this problem are asymptotically equivalent under the null hypothesis, hence also under sequences of contiguous alternatives. We show that this equivalence does not survive asymptotic scenarios where $λ_{n1}/λ_{n2}=1+O(r_n)$ with $r_n=O(1/\sqrt{n})$. For such scenarios, the Le Cam optimal test still asymptotically meets the nominal level constraint, whereas the LRT severely overrejects the null hypothesis. Consequently, the former test should be favored over the latter one whenever the two largest sample eigenvalues are close to each other. By relying on the Le Cam's asymptotic theory of statistical experiments, we study the non-null and optimality properties of the Le Cam optimal test in the aforementioned asymptotic scenarios and show that the null robustness of this test is not obtained at the expense of power. Our asymptotic investigation is extensive in the sense that it allows $r_n$ to converge to zero at an arbitrary rate. While we restrict to single-spiked spectra of the form $λ_{n1}>λ_{n2}=\ldots=λ_{np}$ to make our results as striking as possible, we extend our results to the more general elliptical case. Finally, we present an illustrative real data example.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Davy Paindaveine, Julien Remy, Thomas Verdebout. 2018-12-30. Testing for Principal Component Directions under Weak Identifiability. https://arxiv.org/abs/1710.05291

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A note on the distribution of the partial correlation coefficient with nonparametrically estimated marginal regressions

There has been much interest in the nonparametric testing of conditional independence in the econometric and statistical literature, but the simplest and potentially most useful method, based on the sample partial correlation, seems to have been overlooked, its distribution only having been investigated in some simple parametric instances. The present note shows that an easy to apply permutation test based on the sample partial correlation with nonparametrically estimated marginal regressions has good large and small sample properties.

math.ST

Instance-Log-Optimality of Portfolio-Based E-Processes and their Sequential Hypothesis Tests

We consider the problem of sequential hypothesis testing using $e$-processes. For a rich class of composite testing problems---which include bounded mean testing, equal mean testing for bounded random tuples, and some key ingredients of two-sample and independence testing as special cases---we show that any $e$-process satisfying a certain sublinear regret bound is asymptotically and almost surely instance-log-optimal for a composite alternative. This is a strong notion of optimality that has not previously been established for the aforementioned problems, and we provide explicit test supermartingales and $e$-processes satisfying this notion in a more general case. Furthermore, we derive matching lower and upper bounds on the expected rejection time in the high-confidence regime for the resulting sequential tests in all of these cases. The proofs of these results make weak, algorithm-agnostic moment assumptions and rely on a proof technique involving the aforementioned regret and a family of numeraire portfolios. Finally, we discuss how all of these theorems hold in a distribution-uniform sense, a notion of log-optimality that is stronger still and seems to be new to the literature.

math.ST

Common Drivers in Sparsely Interacting Hawkes Processes

We study a multivariate Hawkes process as a model for time-continuous relational event networks. The model does not assume the network to be known, it includes covariates, and it allows for both common drivers, parameters common to all the actors in the network, and also local parameters specific for each actor. We derive rates of convergence for all of the model parameters when both the number of actors and the time horizon tends to infinity. To prevent an exploding network, sparseness is assumed. We also discuss numerical aspects.

math.ST