arXiv ScienceSearch

arXiv subjects

Roman Vershynin

Publications and source records attributed to Roman Vershynin.

At least 19 recordsLinked to original sources

On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note

We prove an elementary bounded-differences inequality for functions of non-isotropic Gaussian vectors. Specifically, if $f$ has bounded coordinate differences and $X\sim\mathcal N(μ,Σ)$, then the resulting concentration bound depends on the condition number $κ(Σ)$. As an application, we answer a question of Simone Bombari concerning the subgaussianity of sign-quantized linear maps $Y=\mathrm{sgn}(Wx)$. In the special case where $f$ is the coordinatewise sign function, an argument was initially suggested to us by Gemini 3.5 Flash without attribution. We subsequently discovered that it closely resembles an earlier argument of Barber and Kolar [Ann. Statist. 46 (2018), Lemma 4.5]. This revision corrects the attribution and documents the episode as an instance of AI-assisted mathematical discovery.

math.PR

Two friendly proofs of the Berry-Esseen theorem

A gem of classical probability, the Berry-Esseen theorem provides a non-asymptotic form of the central limit theorem. This note gives a friendly and intuitive exposition of two different proofs of the Berry-Esseen theorem for nonidentically distributed random variables: the classical Fourier-analytic proof and a proof by Stein's method following E. Bolthausen. Both proofs are self-contained; the reader may read either proof without reading the other. The exposition is suitable for a basic graduate course in probability.

math.PR

Random sets are close to low-discrepancy sets

We show that a random sample from an arbitrary probability measure on $\mathbb{R}^d$ is close to a low-discrepancy point set. Namely, after moving only a small fraction of the sample points in expectation, one obtains an $n$-point set with star discrepancy $\operatorname{polylog}(n)/n$ with respect to the original measure.

math.PR

Discrepancy and Fisher information

We give an online algorithm that keeps a symmetric random walk inside a convex body by discarding some of its steps. The expected number of discarded steps is controlled by a Fisher-information-type quantity associated with the body. For the cube, this gives a dimension-free bound: a walk with unit Euclidean steps can be kept bounded in all coordinates while discarding only a small constant fraction of the steps on average.

math.PR

Can we spot a fake?

The problem of detecting fake data inspires the following seemingly simple mathematical question. Sample a data point $X$ from the standard normal distribution in $\mathbb{R}^n$. An adversary observes $X$ and corrupts it by adding a vector $rt$, where they can choose any vector $t$ from a fixed set $T$ of the adversary's ``tricks'', and where $r>0$ is a fixed radius. The adversary's choice of $t=t(X)$ may depend on the true data $X$. The adversary wants to hide the corruption by making the fake data $X+rt$ statistically indistinguishable from the real data $X$. What is the largest radius $r=r(T)$ for which the adversary can create an undetectable fake? We show that for highly symmetric sets $T$, the detectability radius $r(T)$ is approximately twice the scaled Gaussian width of $T$. The upper bound actually holds for arbitrary sets $T$ and generalizes to arbitrary, non-Gaussian distributions of real data $X$. The lower bound may fail for not highly symmetric $T$, but we conjecture that this problem can be solved by considering the focused version of the Gaussian width of $T$, which focuses on the most important directions of $T$.

math.ST

Thinning to improve two-sample discrepancy

The discrepancy between two independent samples \(X_1,\dots,X_n\) and \(Y_1,\dots,Y_n\) drawn from the same distribution on $\mathbb{R}^d$ typically has order \(O(\sqrt{n})\) even in one dimension. We give a simple online algorithm that reduces the discrepancy to \(O(\log^{2d} n)\) by discarding a small fraction of the points.

math.PR

Concentration inequalities for random tensors

We show how to extend several basic concentration inequalities for simple random tensors $X = x_1 \otimes \cdots \otimes x_d$ where all $x_k$ are independent random vectors in $\mathbb{R}^n$ with independent coefficients. The new results have optimal dependence on the dimension $n$ and the degree $d$. As an application, we show that random tensors are well conditioned: $(1-o(1)) n^d$ independent copies of the simple random tensor $X \in \mathbb{R}^{n^d}$ are far from being linearly dependent with high probability. We prove this fact for any degree $d = o(\sqrt{n/\log n})$ and conjecture that it is true for any $d = O(n)$.

math.PR

Covering the hypercube, the uncertainty principle, and an interpolation formula

We show that the minimal number of skewed hyperplanes that cover the hypercube $\{0,1\}^{n}$ is at least $\frac{n}{2}+1$, and there are infinitely many $n$'s when the hypercube can be covered with $n-\log_{2}(n)+1$ skewed hyperplanes. The minimal covering problems are closely related to uncertainty principle on the hypercube, where we also obtain an interpolation formula for multilinear polynomials on $\mathbb{R}^{n}$ of degree less than $\lfloor n/m \rfloor$ by showing that its coefficients corresponding to the largest monomials can be represented as a linear combination of values of the polynomial over the points $\{0,1\}^{n}$ whose hamming weights are divisible by $m$.

math.CO

Improving discrepancy by moving a few points

We show how to improve the discrepancy of an iid sample by moving only a few points. Specifically, modifying \( O(m) \) sample points on average reduces the Kolmogorov-Smirnov distance to the population distribution to \(1/m\).

math.ST

Random matrices acting on sets: Independent columns

We study random matrices with independent subgaussian columns. Assuming each column has a fixed Euclidean norm, we establish conditions under which such matrices act as near-isometries when restricted to a given subset of their domain. We show that, with high probability, the maximum distortion caused by such a matrix is proportional to the Gaussian complexity of the subset, scaled by the subgaussian norm of the matrix columns. This linear dependence on the subgaussian norm is a new phenomenon, as random matrices with independent rows or independent entries typically exhibit superlinear dependence. As a consequence, normalizing the columns of random sparse matrices leads to stronger embedding guarantees.

math.PR

LLM Watermarking Using Mixtures and Statistical-to-Computational Gaps

Given a text, can we determine whether it was generated by a large language model (LLM) or by a human? A widely studied approach to this problem is watermarking. We propose an undetectable and elementary watermarking scheme in the closed setting. Also, in the harder open setting, where the adversary has access to most of the model, we propose an unremovable watermarking scheme.

cs.CR

Are most Boolean functions determined by low frequencies?

We ask whether most Boolean functions are determined by their low frequencies. We show a partial result: for almost every function $f: \{-1,1\}^p \to \{-1,1\}$ there exists a function $f': \{-1,1\}^p \to (-1,1)$ that has the same frequencies as $f$ up to dimension $(1/2-o(1))p$.

math.CO

Differentially Private Low-dimensional Synthetic Data from High-dimensional Datasets

Differentially private synthetic data provide a powerful mechanism to enable data analysis while protecting sensitive information about individuals. However, when the data lie in a high-dimensional space, the accuracy of the synthetic data suffers from the curse of dimensionality. In this paper, we propose a differentially private algorithm to generate low-dimensional synthetic data efficiently from a high-dimensional dataset with a utility guarantee with respect to the Wasserstein distance. A key step of our algorithm is a private principal component analysis (PCA) procedure with a near-optimal accuracy bound that circumvents the curse of dimensionality. Unlike the standard perturbation analysis, our analysis of private PCA works without assuming the spectral gap for the covariance matrix.

cs.LG

Hamiltonicity of Sparse Pseudorandom Graphs

We show that every $(n,d,λ)$-graph contains a Hamilton cycle for sufficiently large $n$, assuming that $d\geq \log^{6}n$ and $λ\leq cd$, where $c=\frac{1}{70000}$. This significantly improves a recent result of Glock, Correia and Sudakov, who obtained a similar result for $d$ that grows polynomially with $n$. The proof is based on a new result regarding the second largest eigenvalue of the adjacency matrix of a subgraph induced by a random subset of vertices, combined with a recent result on connecting designated pairs of vertices by vertex-disjoint paths in $(n,d,λ)$-graphs. We believe that the former result is of independent interest and will have further applications.

math.CO

Jackson's inequality on the hypercube

We investigate the best constant $J(n,d)$ such that Jackson's inequality \[ \inf_{\mathrm{deg}(g) \leq d} \|f - g\|_{\infty} \leq J(n,d) \, s(f), \] holds for all functions $f$ on the hypercube $\{0,1\}^n$, where $s(f)$ denotes the sensitivity of $f$. We show that the quantity $J(n, 0.499n)$ is bounded below by an absolute positive constant, independent of $n$. This complements Wagner's theorem, which establishes that $J(n,d)\leq 1 $. As a first application we show that reverse Bernstein inequality fails in the tail space $L^{1}_{\geq 0.499n}$ improving over previously known counterexamples in $L^{1}_{\geq C \log \log (n)}$. As a second application, we show that there exists a function $f : \{0,1\}^n \to [-1,1]$ whose sensitivity $s(f)$ remains constant, independent of $n$, while the approximate degree grows linearly with $n$. This result implies that the sensitivity theorem $s(f) \geq Ω(\mathrm{deg}(f)^C)$ fails in the strongest sense for bounded real-valued functions even when $\mathrm{deg}(f)$ is relaxed to the approximate degree. We also show that in the regime $d = (1 - δ)n$, the bound \[ J(n,d) \leq C \min\{δ, \max\{δ^2, n^{-2/3}\}\} \] holds. Moreover, when restricted to symmetric real-valued functions, we obtain $J_{\mathrm{symmetric}}(n,d) \leq C/d$ and the decay $1/d$ is sharp. Finally, we present results for a subspace approximation problem: we show that there exists a subspace $E$ of dimension $2^{n-1}$ such that $\inf_{g \in E} \|f - g\|_{\infty} \leq s(f)/n$ holds for all $f$.

math.FA

Online Differentially Private Synthetic Data Generation

We present a polynomial-time algorithm for online differentially private synthetic data generation. For a data stream within the hypercube $[0,1]^d$ and an infinite time horizon, we develop an online algorithm that generates a differentially private synthetic dataset at each time $t$. This algorithm achieves a near-optimal accuracy bound of $O(\log(t)t^{-1/d})$ for $d\geq 2$ and $O(\log^{4.5}(t)t^{-1})$ for $d=1$ in the 1-Wasserstein distance. This result extends the previous work on the continual release model for counting queries to Lipschitz queries. Compared to the offline case, where the entire dataset is available at once, our approach requires only an extra polylog factor in the accuracy bound.

math.ST

Differentially Private Synthetic High-dimensional Tabular Stream

While differentially private synthetic data generation has been explored extensively in the literature, how to update this data in the future if the underlying private data changes is much less understood. We propose an algorithmic framework for streaming data that generates multiple synthetic datasets over time, tracking changes in the underlying private data. Our algorithm satisfies differential privacy for the entire input stream (continual differential privacy) and can be used for high-dimensional tabular data. Furthermore, we show the utility of our method via experiments on real-world datasets. The proposed algorithm builds upon a popular select, measure, fit, and iterate paradigm (used by offline synthetic data generation algorithms) and private counters for streams.

cs.CR