arXiv ScienceSearch

arXiv subjects

Junren Chen

Publications and source records attributed to Junren Chen.

3 recordsLinked to original sources

Instance Optimal Sparse Recovery from Nonlinear Observations: A Unified Framework

This paper develops a unified framework for instance optimal sparse recovery from nonlinear observations. The main ingredient is a signal-dependent restricted approximate invertibility condition (RAIC) of some gradient, which leads to the instance optimality of iterative hard thresholding. Under Gaussian designs, we apply the proposed framework to phaseless, one-bit, and ReLU measurements, which correspond to the problems of sparse phase retrieval, one-bit compressed sensing, and sparse ReLU regression, respectively. For sparse phase retrieval, we propose a variant of thresholded amplitude flow and show its instance optimality under $O(s^3)$ measurements (up to logarithmic factors), where $s$ is the sparsity level. To our best knowledge, this is the first instance optimal efficient algorithm for sparse phase retrieval and complements Gao, Wang and Xu (2016) that achieved this via a computationally intractable program. In one-bit compressed sensing, we establish the instance optimality of normalized binary iterative hard thresholding and strengthen the recent result of Matsumoto and Mazumdar (2024). In sparse ReLU regression, it is shown that a slight variant of the algorithm in Soltanolkotabi (2017) is instance optimal. Moreover, $(\ell_2,\ell_2)$ non-uniform instance optimal guarantees are obtained for these problems. The analysis is built upon a number of high-dimensional concentration bounds, including bounds on restricted eigenvalues and a novel instance-dependent hyperplane tessellation result.

cs.IT

On the Subsample Size of Quantile-Based Randomized Kaczmarz

Quantile-based randomized Kaczmarz (QRK) was recently introduced to efficiently solve sparsely corrupted linear systems $\mathbf{A} \mathbf{x}^*+\mathbfε = \mathbf{b}$ [SIAM J. Matrix Anal. Appl., 43(2), 605-637], where $\mathbf{A}\in \mathbb{R}^{m\times n}$ and $\mathbfε$ is an arbitrary $(βm)$-sparse corruption. However, all existing theoretical guarantees for QRK require quantiles to be computed using all $m$ samples (or a subsample of the same order), thus negating the computational advantage of Kaczmarz-type methods. This paper overcomes the bottleneck. We analyze a subsampling QRK, which computes quantiles from $D$ uniformly chosen samples at each iteration. Under some standard scaling assumptions on the coefficient matrix, we show that QRK with subsample size $D\ge\frac{C\log (T)}{\log(1/β)}$ linearly converges over the first $T$ iterations with high probability, where $C$ is some absolute constant. This subsample size is a substantial reduction from $O(m)$ in prior results. For instance, it translates into $O(\log(n))$ even if an approximation error of $\exp(-n^2)$ is desired. Intriguingly, our subsample size is also tight up to a multiplicative constant: if $D\le \frac{c\log(T)}{\log(1/β)}$ for some constant $c$, the error of the $T$-th iterate could be arbitrarily large with high probability. Numerical results are provided to corroborate our theory.

math.NA

Quantile Randomized Kaczmarz for Streaming Linear Systems with Massart Noise

Quantile randomized Kaczmarz (QRK) has proven to be an efficient solver for corrupted linear systems and has received much attention. It was recently shown by Cai et al. (SIAM J. Matrix Anal. Appl. 47(2):802-823, 2026) that using $O(\log T/\log(1/β))$ samples for computing the quantile is necessary and sufficient for QRK to converge linearly over $T$ iterations when solving linear systems with a $β$-fraction of arbitrary corruptions, as long as $β$ is small enough. However, it remains unclear how large the corruption level $β$ can be, and how to compute the required subsample size $D$ explicitly, without hidden constants. This paper studies streaming linear systems with Massart noise via QRK using an order-optimal batch size $D=O(\log T)$ in each update. The independence of samples from previous iterations in the streaming setting enables a sharper analysis, yielding explicit, computable bounds on both the tolerable corruption level and the required subsample size. In particular, we establish linear convergence for corruption levels of up to approximately 7%. We also discuss how the constants improve under oblivious noise.

math.NA