arXiv ScienceSearch

arXiv subjects

Reese Pathak

Publications and source records attributed to Reese Pathak.

17 recordsLinked to original sources

Beyond Modern Asymptotics for Log-Likelihood Ratios in Logistic Regression

We characterize the finite sample behavior of the log-likelihood ratio statistic in binary logistic regression, uniformly over both the design and the target parameter. For $n\geq d\geq 3$, we determine, up to universal constants, its worst case $(1-\delta)$ quantile over all fixed collections of design vectors and all target parameters: \[ d\log\left(\frac{e n}{d}\right)+\log\left(\frac{1}{\delta}\right). \] This is a nonasymptotic analogue of the Wilks $\chi^2_d$ phenomenon and requires no regularity assumptions on the design. The low dimensional cases exhibit unusual behavior. The worst case quantile in dimension $d=2$ is sharply of order \[ \log\log\log n+\log\left(\frac{1}{\delta}\right). \] The worst case quantile in dimension $d=1$ is of order $\log(1/\delta)$, with no dependence on $n$. Finally, i.i.d. Gaussian design vectors recover the classical Wilks scale. In the regime $n\gtrsim d+\log(1/\delta)$, we prove the sharp bound \[ d+\log\left(\frac{1}{\delta}\right). \] Unlike existing asymptotic results, our bounds are uniform over the target parameter, which may depend on $n$, $d$, and $\delta$.

math.ST

Optimal mean width and metric entropy estimates for convex bodies

We show that there exists a constant $C > 0$ such that for any $n \geq 1$ and any convex body $K \subset \mathbf{R}^n$, \[ 1 \leq \inf_{T \in \mathrm{GL}(n)} \, \frac{M^\ast(TK)}{\mathrm{vr}(TK)} \leq C\sqrt{\log(\mathrm{e} n)}, \] where $M^\ast$ denotes the spherical mean width and $\mathrm{vr}(\cdot)$ denotes the volume radius. The righthand side is attained, up to universal constants, by the crosspolytope and the regular $n$-simplex. Analogously, we show that, up to universal constants, the logarithm of the Euclidean covering number is maximized over convex bodies $K \subset \mathbf{R}^n$ by the simplex and crosspolytope. Our proof makes use of Eldan's stochastic localization.

math.MG

Minimum Norm Interpolation via The Local Theory of Banach Spaces: The Role of Gaussianity

We study minimum-norm interpolation (MNI) in overparameterized linear regression with isotropic Gaussian covariates, in settings where the MNI has no closed-form formula. Whereas most prior work relied on Gaussian comparison tools such as the convex Gaussian min--max theorem (CGMT), our approach uses tools from high-dimensional geometry and probability. First, when the norm is in isotropic position, we obtain an ``offset'' bound that controls the amount by which the MNI shrinks the ground truth. Second, we show that the ``intrinsic'' variance of the $\ell_1$-MNI is at most $O(\tfrac{1}{n\log(d/n)^2})$, using a variant of Talagrand's $L_1$--$L_2$ inequality due to Cordero-Erausquin and Ledoux [2012], together with a classical result of Gluskin [1988]. We recover the sharp mean-squared error (MSE) bound for the $\ell_1$-MNI obtained by Wang et al. [2022], using the work of Fleury [2012] on the symmetric Gaussian polytope, which is defined via \[ P_{n,d} := \mathrm{conv}\{\pm X_i\}_{i=1}^{d} \text{ where } X_i \overset{\mathrm{i.i.d.}}{\sim} N(0,\mathrm{I}_{n \times n}), \] rather than CGMT. Our methods also imply improvements on previous results in high-dimensional geometry that may be of independent interest. First, we show that with overwhelming probability, the ratio between the isotropic constant of $P_{n,d}$ and that of the Euclidean ball in $\mathbb{R}^n$ is at most $1+O((\log(d/n))^{-2})$, improving a result of Klartag and Kozma [2009]. We also establish a refined weighted thin-shell estimate on $P_{n,d}$, and provide an elementary proof of the main theorem of Fleury [2012].

math.ST

On the metric projection onto a convex set: reverse H\"older inequalities and upper bounds

We study the $L^p(\mu)$-norm of the metric projection onto a closed, convex set $C \subset \mathbf{R}^n$ when $\mu$ is the uniform measure on the sphere or the standard Gaussian measure on $\mathbf{R}^n$. Up to universal constants, we determine the optimal reverse H\"older inequalities (i.e., $L^q-L^p$ estimates for $q > p$) for both settings and for all $1 \leq p < q \leq \infty$. The optimal constants in these inequalities depend polynomially on the dimension $n$. We establish upper bounds for the expected norm of the metric projection for a wide class of probability measures. Our inequalities improve and extend previous results of S. Chatterjee.

math.MG

A remark on the majorizing measures theorem for general processes

We show that the lower bound in the majorizing measures theorem holds for a large class of random vectors. Specifically, suppose $X \sim \mu$ is a centered random vector in $\mathbf{R}^n$ with \[ C_{\mathrm{KL}}(\mu) = \sup_{\substack{\theta \neq \eta \\ \theta, \eta \in \mathbf{R}^n}} \frac{\mathrm{KL}(\mu_\theta \| \mu_\eta)}{\|\theta - \eta\|_2^2} < \infty, \] where $\mu_\theta$ denotes the law of the translate $\theta + X$. Then, for every nonempty, bounded $T \subset \mathbf{R}^n$, \[ \sqrt{C_{\mathrm{KL}}(\mu)}\, \mathbf{E}_\mu \Big[\sup_{t \in T} \, \langle X, t \rangle \Big] \gtrsim \gamma_2(T), \] where the righthand side denotes Talagrand's generic chaining functional. This result recovers, as a special case, the lower bound in the majorizing measures theorem for centered Gaussian processes. Our argument critically relies on the rate-distortion integral, recently introduced by J. Liu.

math.PR

Gaussian Width of Convex Sets via Integral Decompositions, Projections, and the Distribution of Intrinsic Volumes

We revisit the problem of bounding the expected supremum of a canonical Gaussian process indexed by a convex set $T \subset \mathbf{R}^d$. We develop two decompositions for the Gaussian width, based on the geometry of the index set. The first decomposition involves metric projections of Gaussians onto rescaled copies of $T$. The second involves fixed points arising from a quadratically penalized variant of the local width. Neither decomposition directly invokes generic chaining constructions. Our results make use of recent work in geometric analysis and Gaussian processes. The work of Chatterjee [Ann. Statist., 2014] characterizes the behavior of the metric projection of a Gaussian random vector onto rescaled copies of $T$ with a variational problem involving localized Gaussian widths. We use these bounds to develop decompositions of the Gaussian width using the local metric structure of $T$. Second, we leverage the work of Vitale [Ann. Probab., 1996] to form a connection between the Wills functional (and hence the intrinsic volumes of $T$) and the first terms that appear in our decompositions. Finally, invoking recent work by Mourtada [J. Eur. Math. Soc., 2025] on the logarithm of the Wills functional, we show that the width is controlled by a single, ''peak index'' of the intrinsic volumes. In the worst case, our bound recovers a local form of the classical Dudley integral.

math.PR

Revisiting mean estimation over $\ell_p$ balls: Is the MLE optimal?

We revisit the problem of mean estimation in the Gaussian sequence model with $\ell_p$ constraints for $p \in [0, \infty]$. We demonstrate two phenomena for the behavior of the maximum likelihood estimator (MLE), which depend on the noise level, the radius of the (quasi)norm constraint, the dimension, and the norm index $p$. First, if $p$ lies between $0$ and $1 + \Theta(\tfrac{1}{\log d})$, inclusive, or if it is greater than or equal to $2$, the MLE is minimax rate-optimal for all noise levels and all constraint radii. On the other hand, for the remaining norm indices -- namely, if $p$ lies between $1 + \Theta(\tfrac{1}{\log d})$ and $2$ -- here is a more striking behavior: the MLE is minimax rate-suboptimal, despite its nonlinearity in the observations, for essentially all noise levels and constraint radii for which nonlinear estimates are necessary for minimax-optimal estimation. Our results imply that when given $n$ independent and identically distributed Gaussian samples, the MLE can be suboptimal by a polynomial factor in the sample size. Our lower bounds are constructive: whenever the MLE is rate-suboptimal, we provide explicit instances on which the MLE provably incurs suboptimal risk. Finally, in the non-convex case -- namely when $p < 1$ -- we develop sharp local Gaussian width bounds, which may be of independent interest.

math.ST

Data-Adaptive Tradeoffs among Multiple Risks in Distribution-Free Prediction

Decision-making pipelines are generally characterized by tradeoffs among various risk functions. It is often desirable to manage such tradeoffs in a data-adaptive manner. As we demonstrate, if this is done naively, state-of-the art uncertainty quantification methods can lead to significant violations of putative risk guarantees. To address this issue, we develop methods that permit valid control of risk when threshold and tradeoff parameters are chosen adaptively. Our methodology supports monotone and nearly-monotone risks, but otherwise makes no distributional assumptions. To illustrate the benefits of our approach, we carry out numerical experiments on synthetic data and the large-scale vision dataset MS-COCO.

stat.ME

On the design-dependent suboptimality of the Lasso

This paper investigates the effect of the design matrix on the ability (or inability) to estimate a sparse parameter in linear regression. More specifically, we characterize the optimal rate of estimation when the smallest singular value of the design matrix is bounded away from zero. In addition to this information-theoretic result, we provide and analyze a procedure which is simultaneously statistically optimal and computationally efficient, based on soft thresholding the ordinary least squares estimator. Most surprisingly, we show that the Lasso estimator -- despite its widespread adoption for sparse linear regression -- is provably minimax rate-suboptimal when the minimum singular value is small. We present a family of design matrices and sparse parameters for which we can guarantee that the Lasso with any choice of regularization parameter -- including those which are data-dependent and randomized -- would fail in the sense that its estimation rate is suboptimal by polynomial factors in the sample size. Our lower bound is strong enough to preclude the statistical optimality of all forms of the Lasso, including its highly popular penalized, norm-constrained, and cross-validated variants.

math.ST

Transformers can optimally learn regression mixture models

Mixture models arise in many regression problems, but most methods have seen limited adoption partly due to these algorithms' highly-tailored and model-specific nature. On the other hand, transformers are flexible, neural sequence models that present the intriguing possibility of providing general-purpose prediction methods, even in this mixture setting. In this work, we investigate the hypothesis that transformers can learn an optimal predictor for mixtures of regressions. We construct a generative process for a mixture of linear regressions for which the decision-theoretic optimal procedure is given by data-driven exponential weights on a finite set of parameters. We observe that transformers achieve low mean-squared error on data generated via this process. By probing the transformer's output at inference time, we also show that transformers typically make predictions that are close to the optimal predictor. Our experiments also demonstrate that transformers can learn mixtures of regressions in a sample-efficient fashion and are somewhat robust to distribution shifts. We complement our experimental observations by proving constructively that the decision-theoretic optimal procedure is indeed implementable by a transformer.

cs.LG

Noisy recovery from random linear observations: Sharp minimax rates under elliptical constraints

Estimation problems with constrained parameter spaces arise in various settings. In many of these problems, the observations available to the statistician can be modelled as arising from the noisy realization of the image of a random linear operator; an important special case is random design regression. We derive sharp rates of estimation for arbitrary compact elliptical parameter sets and demonstrate how they depend on the distribution of the random linear operator. Our main result is a functional that characterizes the minimax rate of estimation in terms of the noise level, the law of the random operator, and elliptical norms that define the error metric and the parameter space. This nonasymptotic result is sharp up to an explicit universal constant, and it becomes asymptotically exact as the radius of the parameter space is allowed to grow. We demonstrate the generality of the result by applying it to both parametric and nonparametric regression problems, including those involving distribution shift or dependent covariates.

math.ST

Optimally tackling covariate shift in RKHS-based nonparametric regression

We study the covariate shift problem in the context of nonparametric regression over a reproducing kernel Hilbert space (RKHS). We focus on two natural families of covariate shift problems defined using the likelihood ratios between the source and target distributions. When the likelihood ratios are uniformly bounded, we prove that the kernel ridge regression (KRR) estimator with a carefully chosen regularization parameter is minimax rate-optimal (up to a log factor) for a large family of RKHSs with regular kernel eigenvalues. Interestingly, KRR does not require full knowledge of likelihood ratios apart from an upper bound on them. In striking contrast to the standard statistical setting without covariate shift, we also demonstrate that a naive estimator, which minimizes the empirical risk over the function class, is strictly sub-optimal under covariate shift as compared to KRR. We then address the larger class of covariate shift problems where the likelihood ratio is possibly unbounded yet has a finite second moment. Here, we propose a reweighted KRR estimator that weights samples based on a careful truncation of the likelihood ratios. Again, we are able to show that this estimator is minimax rate-optimal, up to logarithmic factors.

math.ST

A new similarity measure for covariate shift with applications to nonparametric regression

We study covariate shift in the context of nonparametric regression. We introduce a new measure of distribution mismatch between the source and target distributions that is based on the integrated ratio of probabilities of balls at a given radius. We use the scaling of this measure with respect to the radius to characterize the minimax rate of estimation over a family of H\"older continuous functions under covariate shift. In comparison to the recently proposed notion of transfer exponent, this measure leads to a sharper rate of convergence and is more fine-grained. We accompany our theory with concrete instances of covariate shift that illustrate this sharp difference.

math.ST

Cluster-and-Conquer: A Framework For Time-Series Forecasting

We propose a three-stage framework for forecasting high-dimensional time-series data. Our method first estimates parameters for each univariate time series. Next, we use these parameters to cluster the time series. These clusters can be viewed as multivariate time series, for which we then compute parameters. The forecasted values of a single time series can depend on the history of other time series in the same cluster, accounting for intra-cluster similarity while minimizing potential noise in predictions by ignoring inter-cluster effects. Our framework -- which we refer to as "cluster-and-conquer" -- is highly general, allowing for any time-series forecasting and clustering method to be used in each step. It is computationally efficient and embarrassingly parallel. We motivate our framework with a theoretical analysis in an idealized mixed linear regression setting, where we provide guarantees on the quality of the estimates. We accompany these guarantees with experimental results that demonstrate the advantages of our framework: when instantiated with simple linear autoregressive models, we are able to achieve state-of-the-art results on several benchmark datasets, sometimes outperforming deep-learning-based approaches.

cs.LG

FedSplit: An algorithmic framework for fast federated optimization

Motivated by federated learning, we consider the hub-and-spoke model of distributed optimization in which a central authority coordinates the computation of a solution among many agents while limiting communication. We first study some past procedures for federated optimization, and show that their fixed points need not correspond to stationary points of the original optimization problem, even in simple convex settings with deterministic updates. In order to remedy these issues, we introduce FedSplit, a class of algorithms based on operator splitting procedures for solving distributed convex minimization with additive structure. We prove that these procedures have the correct fixed points, corresponding to optima of the original optimization problem, and we characterize their convergence rates under different settings. Our theory shows that these methods are provably robust to inexact computation of intermediate local quantities. We complement our theory with some simple experiments that demonstrate the benefits of our methods in practice.

cs.LG

On Identifying and Mitigating Bias in the Estimation of the COVID-19 Case Fatality Rate

The relative case fatality rates (CFRs) between groups and countries are key measures of relative risk that guide policy decisions regarding scarce medical resource allocation during the ongoing COVID-19 pandemic. In the middle of an active outbreak when surveillance data is the primary source of information, estimating these quantities involves compensating for competing biases in time series of deaths, cases, and recoveries. These include time- and severity- dependent reporting of cases as well as time lags in observed patient outcomes. In the context of COVID-19 CFR estimation, we survey such biases and their potential significance. Further, we analyze theoretically the effect of certain biases, like preferential reporting of fatal cases, on naive estimators of CFR. We provide a partially corrected estimator of these naive estimates that accounts for time lag and imperfect reporting of deaths and recoveries. We show that collection of randomized data by testing the contacts of infectious individuals regardless of the presence of symptoms would mitigate bias by limiting the covariance between diagnosis and death. Our analysis is supplemented by theoretical and numerical results and a simple and fast open-source codebase at https://github.com/aangelopoulos/cfr-covid-19 .

q-bio.QM

Weighted matrix completion from non-random, non-uniform sampling patterns

We study the matrix completion problem when the observation pattern is deterministic and possibly non-uniform. We propose a simple and efficient debiased projection scheme for recovery from noisy observations and analyze the error under a suitable weighted metric. We introduce a simple function of the weight matrix and the sampling pattern that governs the accuracy of the recovered matrix. We derive theoretical guarantees that upper bound the recovery error and nearly matching lower bounds that showcase optimality in several regimes. Our numerical experiments demonstrate the computational efficiency and accuracy of our approach, and show that debiasing is essential when using non-uniform sampling patterns.

cs.IT