arXiv ScienceSearch

arXiv subjects

Bodhisattva Sen

Publications and source records attributed to Bodhisattva Sen.

At least 19 recordsLinked to original sources

Constrained Denoising, Empirical Bayes, and Optimal Transport

In latent variables models, two important goals are denoising and deconvolution: denoising aims to estimate the latent variables, whereas deconvolution aims to estimate the distribution of the latent variables. As has been recognized in the literature over the last century, these two goals are fundamentally in tension, since denoising yields a poor estimate of the distribution of the latent variables due to shrinkage, and deconvolution yields a distribution-valued estimate that carries no unit-specific information. In this paper, we provide a systematic study of denoisers, and empirical Bayes approximations thereof, which attain optimal denoising error subject to the constraint that the distribution of the denoised data matches, in some sense, the distribution of the latent variables. Our insight is that optimal transport allows practitioners to navigate the tension between denoising and deconvolution. More precisely, we propose a modular methodology that combines any suitable unconstrained empirical Bayes denoiser (arising, e.g., via $F$-modeling, $G$-modeling, conjugate-parametric models) with any suitable information about the distribution of the latent variables (e.g., its moments, support, or an approximation of the entire distribution via deconvolution) into a single denoised data set. We prove explicit rates of convergence for our proposed methodologies, and we apply the resulting methods in applications in astronomy, baseball analytics, and marketing.

stat.ME

On the Wasserstein alignment problem

Suppose we are given two metric spaces and a family of continuous transformations from one to the other. Given a probability distribution on each of these two spaces -- namely the source and the target measures -- the Wasserstein alignment problem seeks the transformation that minimizes the optimal transport cost between the pushforward of the source distribution and the target distribution, ensuring the closest possible alignment in a probabilistic sense. Examples of interest include two distributions on two Euclidean spaces $\mathbb{R}^n$ and $\mathbb{R}^d$, and we want a spatial embedding of the $n$-dimensional source measure in $\mathbb{R}^d$ that is closest in some Wasserstein metric to the target distribution on $\mathbb{R}^d$. Similar data alignment problems also commonly arise in shape analysis and computer vision. In this paper, we show that this nonconvex optimal transport projection problem admits a convex Kantorovich-type dual that exploits statistical independence. This allows us to characterize the set of projections and devise a linear programming algorithm. For certain special examples, such as orthogonal transformations on Euclidean spaces of unequal dimensions and the $2$-Wasserstein cost, we characterize the covariance of the optimal projections. Our results also cover the generalization when we penalize each transformation by a function. An example is the inner product Gromov--Wasserstein distance minimization problem which has recently gained popularity.

math.PR

A Mean Field Approach to Empirical Bayes Estimation in High-dimensional Linear Regression

We study empirical Bayes estimation in high-dimensional linear regression. To facilitate computationally efficient estimation of the underlying prior, we adopt a variational empirical Bayes approach, introduced originally in Carbonetto and Stephens (2012) and Kim et al. (2022). We establish asymptotic consistency of the nonparametric maximum likelihood estimator (NPMLE) and its (computable) naive mean field variational surrogate under mild assumptions on the design and the prior. Assuming, in addition, that the naive mean field approximation has a dominant optimizer, we develop a computationally efficient approximation to the oracle posterior distribution, and establish its accuracy under the 1-Wasserstein metric. This enables computationally feasible Bayesian inference; e.g., construction of posterior credible intervals with an average coverage guarantee, Bayes optimal estimation for the regression coefficients, estimation of the proportion of non-nulls, etc. Our analysis covers both deterministic and random designs, and accommodates correlations among the features. To the best of our knowledge, this provides the first rigorous nonparametric empirical Bayes method in a high-dimensional regression setting without sparsity.

math.ST

Empirical Bayes Estimation and Inference via Smooth Nonparametric Maximum Likelihood

The empirical Bayes $g$-modeling approach based on the nonparametric maximum likelihood estimator (NPMLE) has been central to large-scale estimation and inference in the normal means problem. However, theoretical guarantees for uncertainty quantification remain scarce. A key obstacle is that the NPMLE is necessarily discrete, which yields discrete posterior credible sets and a slow logarithmic deconvolution rate. We address both limitations by introducing a hierarchical Gaussian smoothing layer that restricts the mixing distribution to a Gaussian location mixture. Our smooth NPMLE inherits the favorable properties of the classical NPMLE: it is computable via convex optimization and achieves nearly parametric denoising performance. Moreover, it achieves a polynomial deconvolution rate that is asymptotically minimax over the corresponding class. Our procedure also leads to estimated smooth posteriors that converge to the true posteriors at a polynomial rate. Further, we characterize marginal coverage sets that are optimal in expected length, construct plug-in estimators of these sets, and establish theoretical guarantees for the estimated sets in terms of both coverage probability and expected length. We also extend the theory to settings with model misspecification and heteroscedastic Gaussian observations, and study identifiability of the proposed hierarchical model.

math.ST

Nonparametric Riemannian Empirical Bayes, and Denoising Measurements on Manifolds

We initiate the study of nonparametric empirical Bayes denoising methods in the setting where both the latent variables and their measurements lie on a compact Riemannian manifold, and where the likelihood is a Riemannian Gaussian distribution. Our starting point is a novel Tweedie-Eddington formula for Riemannian Gaussian mixture models which identifies a certain surrogate oracle denoiser in terms of the marginal distribution of the measurements; it avoids the explicit computation of the posterior Fréchet mean (as required by the Bayes denoiser) via a first-order approximation, hence we refer to it as the "tangential" Bayes denoiser. We show that this surrogate oracle achieves nearly the Bayes risk in a low-noise regime, we construct a fully data-driven approximation of it using the spectral theory of the Laplace-Beltrami operator, and we establish finite-sample rates of convergence for the distance between the the surrogate oracle and its approximation. Contrasting the nearly-parametric rates from the Euclidean setting, the rates in the Riemannian setting are slower due to the singularities of the Riemannian Gaussian density at the cut locus of its Fréchet mean; in the special case of the circle we establish matching lower bounds which show that our proposed denoiser is minimax-optimal, and that the denoising problem exhibits a genuinely nonparametric rate of convergence. Lastly, we implement our methodology in two scientific applications: in astronomy, the sphere-valued problem of denoising the locations of gamma ray bursts; in structural biology, the torus-valued problem of denoising pairs of torsion angles of adjacent amino acids in a protein (i.e., the Ramachandran plot).

math.ST

Robustness and Efficiency of Rosenbaum's Rank-based Estimator in Randomized Trials: A Design-based Perspective

Mean-based estimators of causal effects in randomized experiments may behave poorly if the potential outcomes have a heavy tail or contain outliers. An alternative estimator proposed by Rosenbaum (1993) estimates a constant additive treatment effect by inverting a randomization test using ranks. We develop a design-based asymptotic theory for this rank-based estimator and study its robustness and efficiency properties. We show that Rosenbaum's estimator is robust against outliers with a breakdown point that uniformly dominates that of any weighted quantile estimator. When pretreatment covariates are available, a regression-adjusted version of Rosenbaum's estimator uses an agnostic linear regression on the covariates and bases inference on the ranks of residuals. Under mild integrability conditions, we show that this estimator is at most 13.6% less efficient, in the worst case, than the commonly used mean-based regression adjustment method proposed by Lin (2013); often outperforming it when the residuals have heavy tails. Moreover, under suitable assumptions, Rosenbaum's regression-adjusted estimator is at least as efficient as the unadjusted one. Finally, we initiate the study of Rosenbaum's estimator when the constant treatment effect assumption may be violated. To analyze the regression-adjusted estimator, we develop local asymptotics of rank statistics under the design-based framework, which may be of independent interest.

stat.ME

Wasserstein-Cramér-Rao Theory of Unbiased Estimation

The quantity of interest in the classical Cramér-Rao theory of unbiased estimation (e.g., the Cramér-Rao lower bound, its exact attainment for exponential families, and asymptotic efficiency of maximum likelihood estimation) is the variance, which represents the instability of an estimator when its value is compared to the value for an independently-sampled data set from the same distribution. In this paper we are interested in a quantity which represents the instability of an estimator when its value is compared to the value for an infinitesimal additive perturbation of the original data set; we refer to this as the "sensitivity" of an estimator. The resulting theory of sensitivity is based on the Wasserstein geometry in the same way that the classical theory of variance is based on the Fisher-Rao (equivalently, Hellinger) geometry, and this insight allows us to determine a collection of results which are analogous to the classical case: a Wasserstein-Cramér-Rao lower bound for the sensitivity of any unbiased estimator, a characterization of models in which there exist unbiased estimators achieving the lower bound exactly, and some concrete results that show that the Wasserstein projection estimator achieves the lower bound asymptotically. We use these results to treat many statistical examples, sometimes revealing new optimality properties for existing estimators and other times revealing entirely new estimators.

math.ST

Estimation of Algebraic Sets: Extending PCA Beyond Linearity

An algebraic set is defined as the zero locus of a system of real polynomial equations. In this paper we address the problem of recovering an unknown algebraic set $\mathcal{A}$ from noisy observations of latent points lying on $\mathcal{A}$ -- a task that extends principal component analysis, which corresponds to the purely linear case. Our procedure consists of three steps: (i) constructing the {\it moment matrix} from the Vandermonde matrix associated with the data set and the degree of the fitted polynomials, (ii) debiasing this moment matrix to remove the noise-induced bias, (iii) extracting its kernel via an eigenvalue decomposition of the debiased moment matrix. These steps yield $n^{-1/2}$-consistent estimators of the coefficients of a set of generators for the ideal of polynomials vanishing on $\mathcal{A}$. To reconstruct $\mathcal{A}$ itself, we propose three complementary strategies: (a) compute the zero set of the fitted polynomials; (b) build a semi-algebraic approximation that encloses $\mathcal{A}$; (c) when structural prior information is available, project the estimated coefficients onto the corresponding constrained space. We prove (nearly) parametric asymptotic error bounds and show that each approach recovers $\mathcal{A}$ under mild regularity conditions.

math.ST

Variational Inference for Latent Variable Models in High Dimensions

Variational inference (VI) is a popular method for approximating intractable posterior distributions in Bayesian inference and probabilistic machine learning. In this paper, we introduce a general framework for quantifying the statistical accuracy of mean-field variational inference (MFVI) for posterior approximation in Bayesian latent variable models with categorical local latent variables (and arbitrary global latent variables). Utilizing our general framework, we capture the exact regime where MFVI 'works' for the celebrated latent Dirichlet allocation model. Focusing on the mixed membership stochastic blockmodel, we show that the vanilla fully factorized MFVI, often used in the literature, is suboptimal. We propose a partially grouped VI algorithm for this model and show that it works, and derive its exact finite-sample performance. We further illustrate that our bounds are tight for both the above models. Our proof techniques, which extend the framework of nonlinear large deviations, open the door for the analysis of MFVI in other latent variable models.

math.ST

Multivariate Distribution-Free Nonparametric Testing: Generalizing Wilcoxon's Tests via Optimal Transport

This paper reviews recent advancements in the application of optimal transport (OT) to multivariate distribution-free nonparametric testing. Inspired by classical rank-based methods, such as Wilcoxon's rank-sum and signed-rank tests, we explore how OT-based ranks and signs generalize these concepts to multivariate settings, while preserving key properties, including distribution-freeness, robustness, and efficiency. Using the framework of asymptotic relative efficiency (ARE), we compare the power of the proposed (generalized Wilcoxon) tests against the Hotelling's $T^2$ test. The ARE lower bounds reveal the Hodges-Lehmann and Chernoff-Savage phenomena in the context of multivariate location testing, underscoring the high power and efficiency of the proposed methods. We also demonstrate how OT-based ranks and signs can be seamlessly integrated with more modern techniques, such as kernel methods, to develop universally consistent, distribution-free tests. Additionally, we present novel results on the construction of consistent and distribution-free kernel-based tests for multivariate symmetry, leveraging OT-based ranks and signs.

stat.ME

Parameter Estimation and Inference in a Continuous Piecewise Linear Regression Model

The estimation of regression parameters in one dimensional broken stick models is a research area of statistics with an extensive literature. We are interested in extending such models by aiming to recover two or more intersecting (hyper)planes in multiple dimensions. In contrast to approaches aiming to recover a given number of piecewise linear components using either a grid search or local smoothing around the change points, we show how to use Nesterov smoothing to obtain a smooth and everywhere differentiable approximation to a piecewise linear regression model with a uniform error bound. The parameters of the smoothed approximation are then efficiently found by minimizing a least squares objective function using a quasi-Newton algorithm. Our main contribution is threefold: We show that the estimates of the Nesterov smoothed approximation of the broken plane model are also $\sqrt{n}$ consistent and asymptotically normal, where $n$ is the number of data points on the two planes. Moreover, we show that as the degree of smoothing goes to zero, the smoothed estimates converge to the unsmoothed estimates and present an algorithm to perform parameter estimation. We conclude by presenting simulation results on simulated data together with some guidance on suitable parameter choices for practical applications.

stat.ME

Distribution-free Measures of Association based on Optimal Transport

In this paper we propose and study a class of nonparametric, yet interpretable measures of association between two random vectors $X$ and $Y$ taking values in $\mathbb{R}^{d_1}$ and $\mathbb{R}^{d_2}$ respectively ($d_1, d_2\ge 1$). These nonparametric measures -- defined using the theory of reproducing kernel Hilbert spaces coupled with optimal transport -- capture the strength of dependence between $X$ and $Y$ and have the property that they are 0 if and only if the variables are independent and 1 if and only if one variable is a measurable function of the other. Further, these population measures can be consistently estimated using the general framework of geometric graphs which include $k$-nearest neighbor graphs and minimum spanning trees. Additionally, these measures can also be readily used to construct an exact finite sample distribution-free test of mutual independence between $X$ and $Y$. In fact, as far as we are aware, these are the only procedures that possess all the above mentioned desirable properties. The correlation coefficient proposed in Dette et al. (2013), Chatterjee (2021), Azadkia and Chatterjee (2021), at the population level, can be seen as a special case of this general class of measures.

math.ST

Empirical partially Bayes multiple testing and compound $χ^2$ decisions

A common task in high-throughput biology is to screen for associations across thousands of units of interest, e.g., genes or proteins. Often, the data for each unit are modeled as Gaussian measurements with unknown mean and variance and are summarized as per-unit sample averages and sample variances. The downstream goal is multiple testing for the means. In this domain, it is routine to "moderate" (that is, to shrink) the sample variances through parametric empirical Bayes methods before computing p-values for the means. Such an approach is asymmetric in that a prior is posited and estimated for the nuisance parameters (variances) but not the primary parameters (means). Our work initiates the formal study of this paradigm, which we term "empirical partially Bayes multiple testing." In this framework, if the prior for the variances were known, one could proceed by computing p-values conditional on the sample variances -- a strategy called partially Bayes inference by Sir David Cox. We show that these conditional p-values satisfy an Eddington/Tweedie-type formula and are approximated at nearly-parametric rates when the prior is estimated by nonparametric maximum likelihood. The estimated p-values can be used with the Benjamini-Hochberg procedure to guarantee asymptotic control of the false discovery rate. Even in the compound setting, wherein the variances are fixed, the approach retains asymptotic type-I error guarantees.

math.ST

A New Perspective On Denoising Based On Optimal Transport

In the standard formulation of the denoising problem, one is given a probabilistic model relating a latent variable $Θ\in Ω\subset \mathbb{R}^m \; (m\ge 1)$ and an observation $Z \in \mathbb{R}^d$ according to: $Z \mid Θ\sim p(\cdot\mid Θ)$ and $Θ\sim G^*$, and the goal is to construct a map to recover the latent variable from the observation. The posterior mean, a natural candidate for estimating $Θ$ from $Z$, attains the minimum Bayes risk (under the squared error loss) but at the expense of over-shrinking the $Z$, and in general may fail to capture the geometric features of the prior distribution $G^*$ (e.g., low dimensionality, discreteness, sparsity, etc.). To rectify these drawbacks, we take a new perspective on this denoising problem that is inspired by optimal transport (OT) theory and use it to study a different, OT-based, denoiser at the population level setting. We rigorously prove that, under general assumptions on the model, this OT-based denoiser is mathematically well-defined and unique, and is closely connected to the solution to a Monge OT problem. We then prove that, under appropriate identifiability assumptions on the model, the OT-based denoiser can be recovered solely from information of the marginal distribution of $Z$ and the posterior mean of the model, after solving a linear relaxation problem over a suitable space of couplings that is reminiscent of standard multimarginal OT problems. In particular, thanks to Tweedie's formula, when the likelihood model $\{ p(\cdot \mid θ) \}_{θ\in Ω}$ is an exponential family of distributions, the OT based-denoiser can be recovered solely from the marginal distribution of $Z$. In general, our family of OT-like relaxations is of interest in its own right and for the denoising problem suggests alternative numerical methods inspired by the rich literature on computational OT.

math.ST

Convex Regression in Multidimensions: Suboptimality of Least Squares Estimators

Under the usual nonparametric regression model with Gaussian errors, Least Squares Estimators (LSEs) over natural subclasses of convex functions are shown to be suboptimal for estimating a $d$-dimensional convex function in squared error loss when the dimension $d$ is 5 or larger. The specific function classes considered include: (i) bounded convex functions supported on a polytope (in random design), (ii) Lipschitz convex functions supported on any convex domain (in random design), (iii) convex functions supported on a polytope (in fixed design). For each of these classes, the risk of the LSE is proved to be of the order $n^{-2/d}$ (up to logarithmic factors) while the minimax risk is $n^{-4/(d+4)}$, when $d \ge 5$. In addition, the first rate of convergence results (worst case and adaptive) for the unrestricted convex LSE are established in fixed-design for polytopal domains for all $d \geq 1$. Some new metric entropy results for convex functions are also proved which are of independent interest.

math.ST

Optimal Confidence Bands for Shape-restricted Regression in Multidimensions

In this paper, we propose and study construction of confidence bands for shape-constrained regression functions when the predictor is multivariate. In particular, we consider the continuous multidimensional white noise model given by $d Y(\mathbf{t}) = n^{1/2} f(\mathbf{t}) \,d\mathbf{t} + d W(\mathbf{t})$, where $Y$ is the observed stochastic process on $[0,1]^d$ ($d\ge 1$), $W$ is the standard Brownian sheet on $[0,1]^d$, and $f$ is the unknown function of interest assumed to belong to a (shape-constrained) function class, e.g., coordinate-wise monotone functions or convex functions. The constructed confidence bands are based on local kernel averaging with bandwidth chosen automatically via a multivariate multiscale statistic. The confidence bands have guaranteed coverage for every $n$ and for every member of the underlying function class. Under monotonicity/convexity constraints on $f$, the proposed confidence bands automatically adapt (in terms of width) to the global and local (Hölder) smoothness and intrinsic dimensionality of the unknown $f$; the bands are also shown to be optimal in a certain sense. These bands have (almost) parametric ($n^{-1/2}$) widths when the underlying function has ``low-complexity'' (e.g., piecewise constant/affine).

math.ST

Multivariate, Heteroscedastic Empirical Bayes via Nonparametric Maximum Likelihood

Multivariate, heteroscedastic errors complicate statistical inference in many large-scale denoising problems. Empirical Bayes is attractive in such settings, but standard parametric approaches rest on assumptions about the form of the prior distribution which can be hard to justify and which introduce unnecessary tuning parameters. We extend the nonparametric maximum likelihood estimator (NPMLE) for Gaussian location mixture densities to allow for multivariate, heteroscedastic errors. NPMLEs estimate an arbitrary prior by solving an infinite-dimensional, convex optimization problem; we show that this convex optimization problem can be tractably approximated by a finite-dimensional version. The empirical Bayes posterior means based on an NPMLE have low regret, meaning they closely target the oracle posterior means one would compute with the true prior in hand. We prove an oracle inequality implying that the empirical Bayes estimator performs at nearly the optimal level (up to logarithmic factors) for denoising without prior knowledge. We provide finite-sample bounds on the average Hellinger accuracy of an NPMLE for estimating the marginal densities of the observations. We also demonstrate the adaptive and nearly-optimal properties of NPMLEs for deconvolution. We apply our method to two denoising problems in astronomy, constructing a fully data-driven color-magnitude diagram of 1.4 million stars in the Milky Way and investigating the distribution of 19 chemical abundance ratios for 27 thousand stars in the red clump. We also apply our method to hierarchical linear models, illustrating the advantages of nonparametric shrinkage of regression coefficients on an education data set and on a microarray data set.

math.ST

Monotone Measure-Preserving Maps in Hilbert Spaces: Existence, Uniqueness, and Stability

The contribution of this work is twofold. The first part deals with a Hilbert-space version of McCann's celebrated result on the existence and uniqueness of monotone measure-preserving maps: given two probability measures $\rm P$ and $\rm Q$ on a separable Hilbert space $\mathcal{H}$ where $\rm P$ does not give mass to "small sets" (namely, Lipschitz hypersurfaces), we show, without imposing any moment assumptions, that there exists a gradient of convex function $\nablaψ$ pushing ${\rm P} $ forward to ${\rm Q}$. In case $\mathcal{H}$ is infinite-dimensional, ${\rm P}$-a.s. uniqueness is not guaranteed, though. If, however, ${\rm Q}$ is boundedly supported (a natural assumption in several statistical applications), then this gradient is ${\rm P}$ a.s. unique. In the second part of the paper, we establish stability results for transport maps in the sense of uniform convergence over compact "regularity sets". As a consequence, we obtain a central limit theorem for the fluctuations of the optimal quadratic transport cost in a separable Hilbert space.

math.PR