arXiv ScienceSearch

arXiv subjects

Mark Rudelson

Publications and source records attributed to Mark Rudelson.

At least 19 recordsLinked to original sources

Well-invertible column subsets of sparse matrices are rare

A random $n\times k$ matrix $S$ is an \emph{$(r,\alpha)$-oblivious subspace injection} (OSI) if $\mathbb{E}\|S^\top x\|_2^2=\|x\|_2^2$ for every $x\in\mathbb{R}^n$, and for every fixed $r$-dimensional subspace $V\subset\mathbb{R}^n$, with probability close to one, one has $\alpha\|x\|_2^2\le\|S^\top x\|_2^2$ for all $x\in V$. In this work, we show that in the regime $r=\Omega(k)$ and $\alpha=\Omega(1)$, and under a mild additional structural assumption, no constant-row-sparsity matrix $S$ is OSI, thereby answering, in a strong form, a question raised by Cama\~no, Epperly, Meyer, and Tropp. We show that the failure of the OSI property for sparse random matrices stems from a general deterministic phenomenon, thereby reducing a probabilistic problem to a non-probabilistic one. This phenomenon is related to the restricted invertibility principle introduced in the seminal work of Bourgain--Tzafriri. Let $(n_k)_{k\in\mathbb{N}}$ be a sequence of integers satisfying $\frac{n_k}{k}\to\infty$. For each $k$, let $S^{(k)}$ be a $n_k\times k$ non-random matrix with $O(1)$ nonzero entries per row, whose nonzero entries have average magnitude $O(1)$, and such that the total number of pairs of rows with supports overlapping at two or more indices is $o({n_k}^2/k)$. We prove that for every constant $\varepsilon>0$, as $k\to\infty$, the overwhelming majority of $k\times \lfloor\varepsilon k\rfloor$ submatrices of $(S^{(k)})^\top$ have the smallest singular value $o(1)$. Thus, the well-invertible submatrices whose existence is guaranteed by the Bourgain--Tzafriri theorem are rare. The proof is itself based on probabilistic tools.

math.PR

Well-Conditioned Oblivious Perturbations in Linear Space

Perturbing a deterministic $n$-dimensional matrix with small Gaussian noise is a cornerstone of smoothed analysis of algorithms [Spielman and Teng, JACM 2004], as it reduces the condition number of the input to $O(n)$, and with it the complexity of many matrix algorithms. However, when deployed algorithmically, these perturbations are expensive due to the cost of generating and storing $n^2$ Gaussian random variables. We propose a perturbation that requires generating and storing $O(n)$ random numbers in $O(\log n)$ bits of precision, and reduces the condition number of any deterministic matrix to $O(n)$, matching Gaussian perturbations. Our result in particular implies a better complexity for the perturbed conjugate gradient algorithm, showing that we can solve an $n\times n$ linear system in linear space to within an arbitrarily small constant backward error using $O(n)$ matrix-vector products. In our construction, we introduce the concept of a pattern matrix, which is a dense deterministic matrix that maps all sparse vectors into dense vectors, and we combine it with a sparse perturbation whose entries are dependent and located in a non-uniform fashion. In order to analyze this construction, we develop new techniques for lower bounding the smallest singular value of a random matrix with dependent entries.

cs.DS

Hardness of approximation of centered convex bodies by polytopes

The distance between convex bodies \(K, L \subseteq \R^n\) is defined as \[ d(K,L)= \inf \left\{ \lambda \ge 1: \ L-x \subseteq T (K-y) \subseteq \lambda (L-x) \right\}, \] where the infimum is taken over all \(x,y \in \R^n\) and all invertible linear operators \(T: \R^n \to \R^n\). If both bodies are centrally symmetric, then the shifts $x$ and $y$ can be chosen to be $0$. In this case, any convex symmetric body $K$ can be approximated by a polytope $P$ with at most $N \in (n, e^{cn})$ vertices so that \[ P \subseteq K \subseteq \lambda P \] where \(\lambda= O \left(\sqrt{\frac{n}{\log N}} \right)\) up to logarithmic factors. We prove that approximating a general centered convex body by a polytope requires a significantly larger number of vertices compared to the symmetric case. More precisely, there exists a convex body \(K \subseteq \R^n\) whose barycenter coincides with the origin, such that any polytope $P$ satisfying \[ P \subseteq K \subseteq c \, \frac{n}{\log N} P \] must have at least \(N\) vertices, provided that \(N \in (Cn^2, e^{cn})\). Moreover, we prove that the same bound holds for approximating a centered convex body with a polytope having $N$ facets instead of $N$ vertices.

math.FA

Optimal Embedding Dimension for Sparse Subspace Embeddings

A random $m\times n$ matrix $S$ is an oblivious subspace embedding (OSE) with parameters $\epsilon>0$, $\delta\in(0,1/3)$ and $d\leq m\leq n$, if for any $d$-dimensional subspace $W\subseteq R^n$, $P\big(\,\forall_{x\in W}\ (1+\epsilon)^{-1}\|x\|\leq\|Sx\|\leq (1+\epsilon)\|x\|\,\big)\geq 1-\delta.$ It is known that the embedding dimension of an OSE must satisfy $m\geq d$, and for any $\theta > 0$, a Gaussian embedding matrix with $m\geq (1+\theta) d$ is an OSE with $\epsilon = O_\theta(1)$. However, such optimal embedding dimension is not known for other embeddings. Of particular interest are sparse OSEs, having $s\ll m$ non-zeros per column, with applications to problems such as least squares regression and low-rank approximation. We show that, given any $\theta > 0$, an $m\times n$ random matrix $S$ with $m\geq (1+\theta)d$ consisting of randomly sparsified $\pm1/\sqrt s$ entries and having $s= O(\log^4(d))$ non-zeros per column, is an oblivious subspace embedding with $\epsilon = O_{\theta}(1)$. Our result addresses the main open question posed by Nelson and Nguyen (FOCS 2013), who conjectured that sparse OSEs can achieve $m=O(d)$ embedding dimension, and it improves on $m=O(d\log(d))$ shown by Cohen (SODA 2016). We use this to construct the first oblivious subspace embedding with $O(d)$ embedding dimension that can be applied faster than current matrix multiplication time, and to obtain an optimal single-pass algorithm for least squares regression. We further extend our results to Leverage Score Sparsification (LESS), which is a recently introduced non-oblivious embedding technique. We use LESS to construct the first subspace embedding with low distortion $\epsilon=o(1)$ and optimal embedding dimension $m=O(d/\epsilon^2)$ that can be applied in current matrix multiplication time.

cs.DS

Approximately Hadamard matrices and Riesz bases in random frames

An $n \times n$ matrix with $\pm 1$ entries which acts on $\mathbb{R}^n$ as a scaled isometry is called Hadamard. Such matrices exist in some, but not all dimensions. Combining number-theoretic and probabilistic tools we construct matrices with $\pm 1$ entries which act as approximate scaled isometries in $\mathbb{R}^n$ for all $n$. More precisely, the matrices we construct have condition numbers bounded by a constant independent of $n$. Using this construction, we establish a phase transition for the probability that a random frame contains a Riesz basis. Namely, we show that a random frame in $\mathbb{R}^n$ formed by $N$ vectors with independent identically distributed coordinates having a non-degenerate symmetric distribution contains many Riesz bases with high probability provided that $N \ge \exp(Cn)$. On the other hand, we prove that if the entries are subgaussian, then a random frame fails to contain a Riesz basis with probability close to $1$ whenever $N \le \exp(cn)$, where $c<C$ are constants depending on the distribution of the entries.

math.PR

A quick estimate for the volume of a polyhedron

Let $P$ be a bounded polyhedron defined as the intersection of the non-negative orthant ${\Bbb R}^n_+$ and an affine subspace of codimension $m$ in ${\Bbb R}^n$. We show that a simple and computationally efficient formula approximates the volume of $P$ within a factor of $\gamma^m$, where $\gamma >0$ is an absolute constant. The formula provides the best known estimate for the volume of transportation polytopes from a wide family.

math.MG

Exact Matching of Random Graphs with Constant Correlation

This paper deals with the problem of graph matching or network alignment for Erd\H{o}s--R\'enyi graphs, which can be viewed as a noisy average-case version of the graph isomorphism problem. Let $G$ and $G'$ be $G(n, p)$ Erd\H{o}s--R\'enyi graphs marginally, identified with their adjacency matrices. Assume that $G$ and $G'$ are correlated such that $\mathbb{E}[G_{ij} G'_{ij}] = p(1-\alpha)$. For a permutation $\pi$ representing a latent matching between the vertices of $G$ and $G'$, denote by $G^\pi$ the graph obtained from permuting the vertices of $G$ by $\pi$. Observing $G^\pi$ and $G'$, we aim to recover the matching $\pi$. In this work, we show that for every $\varepsilon \in (0,1]$, there is $n_0>0$ depending on $\varepsilon$ and absolute constants $\alpha_0, R > 0$ with the following property. Let $n \ge n_0$, $(1+\varepsilon) \log n \le np \le n^{\frac{1}{R \log \log n}}$, and $0 < \alpha < \min(\alpha_0,\varepsilon/4)$. There is a polynomial-time algorithm $F$ such that $\mathbb{P}\{F(G^\pi,G')=\pi\}=1-o(1)$. This is the first polynomial-time algorithm that recovers the exact matching between vertices of correlated Erd\H{o}s--R\'enyi graphs with constant correlation with high probability. The algorithm is based on comparison of partition trees associated with the graph vertices.

math.ST

When a system of real quadratic equations has a solution

We provide a sufficient condition for solvability of a system of real quadratic equations $p_i(x)=y_i$, $i=1, \ldots, m$, where $p_i: {\mathbb R}^n \longrightarrow {\mathbb R}$ are quadratic forms. By solving a positive semidefinite program, one can reduce it to another system of the type $q_i(x)=\alpha_i$, $i=1, \ldots, m$, where $q_i: {\mathbb R}^n \longrightarrow {\mathbb R}$ are quadratic forms and $\alpha_i=\mathrm{tr\ } q_i$. We prove that the latter system has solution $x \in {\mathbb R}^n$ if for some (equivalently, for any) orthonormal basis $A_1,\ldots, A_m$ in the space spanned by the matrices of the forms $q_i$, the operator norm of $A_1^2 + \ldots + A_m^2$ does not exceed $\eta/m$ for some absolute constant $\eta > 0$. The condition can be checked in polynomial time and is satisfied, for example, for random $q_i$ provided $m \leq \gamma \sqrt{n}$ for an absolute constant $\gamma >0$. We prove a similar sufficient condition for a system of homogeneous quadratic equations to have a non-trivial solution. While the condition we obtain is of an algebraic nature, the proof relies on analytic tools including Fourier analysis and measure concentration.

math.OC

Random Graph Matching with Improved Noise Robustness

Graph matching, also known as network alignment, refers to finding a bijection between the vertex sets of two given graphs so as to maximally align their edges. This fundamental computational problem arises frequently in multiple fields such as computer vision and biology. Recently, there has been a plethora of work studying efficient algorithms for graph matching under probabilistic models. In this work, we propose a new algorithm for graph matching: Our algorithm associates each vertex with a signature vector using a multistage procedure and then matches a pair of vertices from the two graphs if their signature vectors are close to each other. We show that, for two Erd\H{o}s--R\'enyi graphs with edge correlation $1-\alpha$, our algorithm recovers the underlying matching exactly with high probability when $\alpha \le 1 / (\log \log n)^C$, where $n$ is the number of vertices in each graph and $C$ denotes a positive universal constant. This improves the condition $\alpha \le 1 / (\log n)^C$ achieved in previous work.

cs.DS

On the volume of non-central sections of a cube

Let $Q_n$ be the cube of side length one centered at the origin in $\mathbb{R}^n$, and let $F$ be an affine $(n-d)$-dimensional subspace of $\mathbb{R}^n$ having distance to the origin less than or equal to $\frac 1 2$, where $0<d<n$. We show that the $(n-d)$-dimensional volume of the section $Q_n \cap F$ is bounded below by a value $c(d)$ depending only on the codimension $d$ but not on the ambient dimension $n$ or a particular subspace $F$. In the case of hyperplanes, $d=1$, we show that $c(1) = \frac{1}{17}$ is a possible choice. We also consider a complex analogue of this problem for a hyperplane section of the polydisc.

math.MG

Size of nodal domains of the eigenvectors of a G(n,p) graph

Consider an eigenvector of the adjacency matrix of a G(n, p) graph. A nodal domain is a connected component of the set of vertices where this eigenvector has a constant sign. It is known that with high probability, there are exactly two nodal domains for each eigenvector corresponding to a non-leading eigenvalue. We prove that with high probability, the sizes of these nodal domains are approximately equal to each other.

math.PR

Restricted Isometry Property under High Correlations

Matrices satisfying the Restricted Isometry Property (RIP) play an important role in the areas of compressed sensing and statistical learning. RIP matrices with optimal parameters are mainly obtained via probabilistic arguments, as explicit constructions seem hard. It is therefore interesting to ask whether a fixed matrix can be incorporated into a construction of restricted isometries. In this paper, we construct a new broad ensemble of random matrices with dependent entries that satisfy the restricted isometry property. Our construction starts with a fixed (deterministic) matrix $X$ satisfying some simple stable rank condition, and we show that the matrix $XR$, where $R$ is a random matrix drawn from various popular probabilistic models (including, subgaussian, sparse, low-randomness, satisfying convex concentration property), satisfies the RIP with high probability. These theorems have various applications in signal recovery, random matrix theory, dimensionality reduction, etc. Additionally, motivated by an application for understanding the effectiveness of word vector embeddings popular in natural language processing and machine learning applications, we investigate the RIP of the matrix $XR^{(l)}$ where $R^{(l)}$ is formed by taking all possible (disregarding order) $l$-way entrywise products of the columns of a random matrix $R$.

cs.LG

Sharp transition of the invertibility of the adjacency matrices of sparse random graphs

We consider three different models of sparse random graphs:~undirected and directed Erd\H{o}s-R\'{e}nyi graphs, and random bipartite graph with an equal number of left and right vertices. For such graphs we show that if the edge connectivity probability $p \in (0,1)$ satisfies $n p \ge \log n + k(n)$ with $k(n) \to \infty$ as $n \to \infty$, then the adjacency matrix is invertible with probability approaching one (here $n$ is the number of vertices in the two former cases and the number of left and right vertices in the latter case). If $np \le \log n -k(n)$ then these matrices are invertible with probability approaching zero, as $n \to \infty$. In the intermediate region, when $np=\log n + k(n)$, for a bounded sequence $k(n) \in \mathbb{R}$, the event $\Omega_0$ that the adjacency matrix has a zero row or a column and its complement both have non-vanishing probability. For such choices of $p$ our results show that conditioned on the event $\Omega_0^c$ the matrices are again invertible with probability tending to one. This shows that the primary reason for the non-invertibility of such matrices is the existence of a zero row or a column. The bounds on the probability of the invertibility of these matrices are a consequence of quantitative lower bounds on their smallest singular values. Combining this with an upper bound on the largest singular value of the centered version of these matrices we show that the (modified) condition number is $O(n^{1+o(1)})$ on the event that there is no zero row or column, with large probability. This matches with von Neumann's prediction about the condition number of random matrices up to a factor of $n^{o(1)}$, for the entire range of $p$.

math.PR

The sparse circular law under minimal assumptions

The circular law asserts that the empirical distribution of eigenvalues of appropriately normalized $n\times n$ matrix with i.i.d. entries converges to the uniform measure on the unit disc as the dimension $n$ grows to infinity. Consider an $n\times n$ matrix $A_n=(\delta_{ij}^{(n)}\xi_{ij}^{(n)})$, where $\xi_{ij}^{(n)}$ are copies of a real random variable of unit variance, variables $\delta_{ij}^{(n)}$ are Bernoulli ($0/1$) with ${\mathbb P}\{\delta_{ij}^{(n)}=1\}=p_n$, and $\delta_{ij}^{(n)}$ and $\xi_{ij}^{(n)}$, $i,j\in[n]$, are jointly independent. In order for the circular law to hold for the sequence $\big(\frac{1}{\sqrt{p_n n}}A_n\big)$, one has to assume that $p_n n\to \infty$. We derive the circular law under this minimal assumption.

math.PR

Delocalization of eigenvectors of random matrices. Lecture notes

Let $x \in S^{n-1}$ be a unit eigenvector of an $n \times n$ random matrix. This vector is delocalized if it is distributed roughly uniformly over the real or complex sphere. This intuitive notion can be quantified in various ways. In these lectures, we will concentrate on the no-gaps delocalization. This type of delocalization means that with high probability, any non-negligible subset of the support of $x$ carries a non-negligible mass. Proving the no-gaps delocalization requires establishing small ball probability bounds for the projections of random vector. Using Fourier transform, we will prove such bounds in a simpler case of a random vector having independent coordinates of a bounded density. This will allow us to derive the no-gaps delocalization for matrices with random entries having a bounded density. In the last section, we will discuss the applications of delocalization to the spectral properties of Erd\H{o}s-R\'enyi random graphs.

math.PR

Restricted Eigenvalue from Stable Rank with Applications to Sparse Linear Regression

High-dimensional settings, where the data dimension ($d$) far exceeds the number of observations ($n$), are common in many statistical and machine learning applications. Methods based on $\ell_1$-relaxation, such as Lasso, are very popular for sparse recovery in these settings. Restricted Eigenvalue (RE) condition is among the weakest, and hence the most general, condition in literature imposed on the Gram matrix that guarantees nice statistical properties for the Lasso estimator. It is natural to ask: what families of matrices satisfy the RE condition? Following a line of work in this area, we construct a new broad ensemble of dependent random design matrices that have an explicit RE bound. Our construction starts with a fixed (deterministic) matrix $X \in \mathbb{R}^{n \times d}$ satisfying a simple stable rank condition, and we show that a matrix drawn from the distribution $X \Phi^\top \Phi$, where $\Phi \in \mathbb{R}^{m \times d}$ is a subgaussian random matrix, with high probability, satisfies the RE condition. This construction allows incorporating a fixed matrix that has an easily {\em verifiable} condition into the design process, and allows for generation of {\em compressed} design matrices that have a lower storage requirement than a standard design matrix. We give two applications of this construction to sparse linear regression problems, including one to a compressed sparse regression setting where the regression algorithm only has access to a compressed representation of a fixed design matrix $X$.

stat.ML

The circular law for sparse non-Hermitian matrices

For a class of sparse random matrices of the form $A_n =(\xi_{i,j}\delta_{i,j})_{i,j=1}^n$, where $\{\xi_{i,j}\}$ are i.i.d.~centered sub-Gaussian random variables of unit variance, and $\{\delta_{i,j}\}$ are i.i.d.~Bernoulli random variables taking value $1$ with probability $p_n$, we prove that the empirical spectral distribution of $A_n/\sqrt{np_n}$ converges weakly to the circular law, in probability, for all $p_n$ such that $p_n=\omega({\log^2n}/{n})$. Additionally if $p_n$ satisfies the inequality $np_n > \exp(c\sqrt{\log n})$ for some constant $c$, then the above convergence is shown to hold almost surely. The key to this is a new bound on the smallest singular value of complex shifts of real valued sparse random matrices. The circular law limit also extends to the adjacency matrix of a directed Erd\H{o}s-R\'{e}nyi graph with edge connectivity probability $p_n$.

math.PR

Errors-in-variables models with dependent measurements

Suppose that we observe $y \in \mathbb{R}^n$ and $X \in \mathbb{R}^{n \times m}$ in the following errors-in-variables model: \begin{eqnarray*} y & = & X_0 \beta^* +\epsilon \\ X & = & X_0 + W, \end{eqnarray*} where $X_0$ is an $n \times m$ design matrix with independent subgaussian row vectors, $\epsilon \in \mathbb{R}^n$ is a noise vector and $W$ is a mean zero $n \times m$ random noise matrix with independent subgaussian column vectors, independent of $X_0$ and $\epsilon$. This model is significantly different from those analyzed in the literature in the sense that we allow the measurement error for each covariate to be a dependent vector across its $n$ observations. Such error structures appear in the science literature when modeling the trial-to-trial fluctuations in response strength shared across a set of neurons. Under sparsity and restrictive eigenvalue type of conditions, we show that one is able to recover a sparse vector $\beta^* \in \mathbb{R}^m$ from the model given a single observation matrix $X$ and the response vector $y$. We establish consistency in estimating $\beta^*$ and obtain the rates of convergence in the $\ell_q$ norm, where $q = 1, 2$. We show error bounds which approach that of the regular Lasso and the Dantzig selector in case the errors in $W$ are tending to 0. We analyze the convergence rates of the gradient descent methods for solving the nonconvex programs and show that the composite gradient descent algorithm is guaranteed to converge at a geometric rate to a neighborhood of the global minimizers: the size of the neighborhood is bounded by the statistical error in the $\ell_2$ norm. Our analysis reveals interesting connections between computational and statistical efficiency and the concentration of measure phenomenon in random matrix theory. We provide simulation evidence illuminating the theoretical predictions.

stat.ML