arXiv ScienceSearch

arXiv subjects

Philip Thompson

Publications and source records attributed to Philip Thompson.

12 recordsLinked to original sources

Robust dimension-free estimation of simple random tensors: optimal guarantees under heavy tails and adversarial contamination

We study robust estimation of simple random tensors of arbitrary order $q\in\mathbb{N}$ under finite-moment assumptions and adversarial contamination. We propose the first robust estimator achieving near-optimal dimension-free statistical rates in this setting. The estimator attains the near-optimal corruption rate whenever $p\ge2q$ moments are finite and continues to provide nontrivial guarantees throughout the weak-moment regime $q\le p\le2q$. Being based on directional trimmed means and minimax aggregation, our estimator is adaptive to $p$ and upper bounds on hypercontractive constants without resorting to interval-intersection procedures. Our analysis extends the trimmed-mean framework underlying recent advances in robust mean and covariance estimation to arbitrary tensor order. In particular, we establish concentration inequalities for higher-order counting and truncated empirical multi-vector product processes. We believe these inequalities could be of independent interest beyond the present application, including algorithmic robust estimation.

math.ST

Outlier-robust additive matrix decomposition

We study least-squares trace regression when the parameter is the sum of a $r$-low-rank matrix and a $s$-sparse matrix and a fraction $\epsilon$ of the labels is corrupted. For subgaussian distributions and feature-dependent noise, we highlight three needed design properties, each one derived from a different process inequality: a "product process inequality", "Chevet's inequality" and a "multiplier process inequality". These properties handle, simultaneously, additive decomposition, label contamination and design-noise interaction. They imply the near-optimality of a tractable estimator with respect to the effective dimensions $d_{eff,r}$ and $d_{eff,s}$ of the low-rank and sparse components, $\epsilon$ and the failure probability $\delta$. The near-optimal rate is $\mathsf{r}(n,d_{eff,r}) + \mathsf{r}(n,d_{eff,s}) + \sqrt{(1+\log(1/\delta))/n} + \epsilon\log(1/\epsilon)$, where $\mathsf{r}(n,d_{eff,r})+\mathsf{r}(n,d_{eff,s})$ is the optimal rate in average with no contamination. Our estimator is adaptive to $(s,r,\epsilon,\delta)$ and, for fixed absolute constant $c>0$, it attains the mentioned rate with probability $1-\delta$ uniformly over all $\delta\ge\exp(-cn)$. Without matrix decomposition, our analysis also entails optimal bounds for a robust estimator adapted to the noise variance. Our estimators are based on "sorted" versions of Huber's loss. We present simulations matching the theory. In particular, it reveals the superiority of "sorted" Huber's losses over the classical Huber's loss.

math.ST

A spectral least-squares-type method for heavy-tailed corrupted regression with unknown covariance \& heterogeneous noise

We revisit heavy-tailed corrupted least-squares linear regression assuming to have a corrupted $n$-sized label-feature sample of at most $\epsilon n$ arbitrary outliers. We wish to estimate a $p$-dimensional parameter $b^*$ given such sample of a label-feature pair $(y,x)$ satisfying $y=\langle x,b^*\rangle+\xi$ with heavy-tailed $(x,\xi)$. We only assume $x$ is $L^4-L^2$ hypercontractive with constant $L>0$ and has covariance matrix $\Sigma$ with minimum eigenvalue $1/\mu^2>0$ and bounded condition number $\kappa>0$. The noise $\xi$ can be arbitrarily dependent on $x$ and nonsymmetric as long as $\xi x$ has finite covariance matrix $\Xi$. We propose a near-optimal computationally tractable estimator, based on the power method, assuming no knowledge on $(\Sigma,\Xi)$ nor the operator norm of $\Xi$. With probability at least $1-\delta$, our proposed estimator attains the statistical rate $\mu^2\Vert\Xi\Vert^{1/2}(\frac{p}{n}+\frac{\log(1/\delta)}{n}+\epsilon)^{1/2}$ and breakdown-point $\epsilon\lesssim\frac{1}{L^4\kappa^2}$, both optimal in the $\ell_2$-norm, assuming the near-optimal minimum sample size $L^4\kappa^2(p\log p + \log(1/\delta))\lesssim n$, up to a log factor. To the best of our knowledge, this is the first computationally tractable algorithm satisfying simultaneously all the mentioned properties. Our estimator is based on a two-stage Multiplicative Weight Update algorithm. The first stage estimates a descent direction $\hat v$ with respect to the (unknown) pre-conditioned inner product $\langle\Sigma(\cdot),\cdot\rangle$. The second stage estimate the descent direction $\Sigma\hat v$ with respect to the (known) inner product $\langle\cdot,\cdot\rangle$, without knowing nor estimating $\Sigma$.

math.ST

Outlier-robust sparse/low-rank least-squares regression and robust matrix completion

We study high-dimensional least-squares regression within a subgaussian statistical learning framework with heterogeneous noise. It includes $s$-sparse and $r$-low-rank least-squares regression when a fraction $\epsilon$ of the labels are adversarially contaminated. We also present a novel theory of trace-regression with matrix decomposition based on a new application of the product process. For these problems, we show novel near-optimal "subgaussian" estimation rates of the form $r(n,d_{e})+\sqrt{\log(1/\delta)/n}+\epsilon\log(1/\epsilon)$, valid with probability at least $1-\delta$. Here, $r(n,d_{e})$ is the optimal uncontaminated rate as a function of the effective dimension $d_{e}$ but independent of the failure probability $\delta$. These rates are valid uniformly on $\delta$, i.e., the estimators' tuning do not depend on $\delta$. Lastly, we consider noisy robust matrix completion with non-uniform sampling. If only the low-rank matrix is of interest, we present a novel near-optimal rate that is independent of the corruption level $a$. Our estimators are tractable and based on a new "sorted" Huber-type loss. No information on $(s,r,\epsilon,a)$ are needed to tune these estimators. Our analysis makes use of novel $\delta$-optimal concentration inequalities for the multiplier and product processes which could be useful elsewhere. For instance, they imply novel sharp oracle inequalities for Lasso and Slope with optimal dependence on $\delta$. Numerical simulations confirm our theoretical predictions. In particular, "sorted" Huber regression can outperform classical Huber regression.

math.ST

Outlier-robust estimation of a sparse linear model using $\ell_1$-penalized Huber's $M$-estimator

We study the problem of estimating a $p$-dimensional $s$-sparse vector in a linear model with Gaussian design and additive noise. In the case where the labels are contaminated by at most $o$ adversarial outliers, we prove that the $\ell_1$-penalized Huber's $M$-estimator based on $n$ samples attains the optimal rate of convergence $(s/n)^{1/2} + (o/n)$, up to a logarithmic factor. For more general design matrices, our results highlight the importance of two properties: the transfer principle and the incoherence property. These properties with suitable constants are shown to yield the optimal rates, up to log-factors, of robust estimation with adversarial contamination.

math.ST

Restricted eigenvalue property for corrupted Gaussian designs

Motivated by the construction of tractable robust estimators via convex relaxations, we present conditions on the sample size which guarantee an augmented notion of Restricted Eigenvalue-type condition for Gaussian designs. Such a notion is suitable for high-dimensional robust inference in a Gaussian linear model and a multivariate Gaussian model when samples are corrupted by outliers either in the response variable or in the design matrix. Our proof technique relies on simultaneous lower and upper bounds of two random bilinear forms with very different behaviors. Such simultaneous bounds are used for balancing the interaction between the parameter vector and the estimated corruption vector as well as for controlling the presence of corruption in the design. Our technique has the advantage of not relying on known bounds of the extreme singular values of the associated Gaussian ensemble nor on the use of mutual incoherence arguments. A relevant consequence of our analysis, compared to prior work, is that a significantly sharper restricted eigenvalue constant can be obtained under weaker assumptions. In particular, the sparsity of the unknown parameter and the number of outliers are allowed to be completely independent of each other.

math.ST

Sample average approximation with heavier tails II: localization in stochastic convex optimization and persistence results for the Lasso

``Localization'' has proven to be a valuable tool in the Statistical Learning literature as it allows sharp risk bounds in terms of the problem geometry. Localized bounds seem to be much less exploited in the Stochastic Optimization literature. In addition, there is an obvious interest in both communities in obtaining risk bounds that require weak moment assumptions or ``heavier-tails''. In this work we use a localization toolbox to derive risk bounds in two specific applications. The first is in portfolio risk minimization with conditional value-at-risk constraints. We consider a setting where, among all assets with high returns, there is a portion of dimension $g$, unknown to the investor, that has significant less risk than the other remaining portion. Our rates for the SAA problem show that ``risk inflation'', caused by a multiplicative factor, affects the statistical rate only via a term proportional to $g$. As the ``normalized risk'' increases, the contribution in the rate from the extrinsic dimension diminishes while the dependence on $g$ is kept fixed. Localization is a key tool to show this property. As a second application of our localization toolbox, we obtain sharp oracle inequalities for least-squares estimators with a Lasso-type constraint under weak moment assumptions. One main consequence of these inequalities is to obtain \emph{persistence}, as posed by Greenshtein and Ritov, with covariates having heavier tails. This gives improvements in prior work of Bartlett, Mendelson and Neeman.

math.OC

On variance reduction for stochastic smooth convex optimization with multiplicative noise

We propose dynamic sampled stochastic approximation (SA) methods for stochastic optimization with a heavy-tailed distribution (with finite 2nd moment). The objective is the sum of a smooth convex function with a convex regularizer. Typically, it is assumed an oracle with an upper bound $\sigma^2$ on its variance (OUBV). Differently, we assume an oracle with \emph{multiplicative noise}. This rarely addressed setup is more aggressive but realistic, where the variance may not be bounded. Our methods achieve optimal iteration complexity and (near) optimal oracle complexity. For the smooth convex class, we use an accelerated SA method a la FISTA which achieves, given tolerance $\epsilon>0$, the optimal iteration complexity of $\mathcal{O}(\epsilon^{-\frac{1}{2}})$ with a near-optimal oracle complexity of $\mathcal{O}(\epsilon^{-2})[\ln(\epsilon^{-\frac{1}{2}})]^2$. This improves upon Ghadimi and Lan [\emph{Math. Program.}, 156:59-99, 2016] where it is assumed an OUBV. For the strongly convex class, our method achieves optimal iteration complexity of $\mathcal{O}(\ln(\epsilon^{-1}))$ and optimal oracle complexity of $\mathcal{O}(\epsilon^{-1})$. This improves upon Byrd et al. [\emph{Math. Program.}, 134:127-155, 2012] where it is assumed an OUBV. In terms of variance, our bounds are local: they depend on variances $\sigma(x^*)^2$ at solutions $x^*$ and the per unit distance multiplicative variance $\sigma^2_L$. For the smooth convex class, there exist policies such that our bounds resemble those obtained if it was assumed an OUBV with $\sigma^2:=\sigma(x^*)^2$. For the strongly convex class such property is obtained exactly if the condition number is estimated or in the limit for better conditioned problems or for larger initial batch sizes. In any case, if it is assumed an OUBV, our bounds are thus much sharper since typically $\max\{\sigma(x^*)^2,\sigma_L^2\}\ll\sigma^2$.

math.OC

Sample average approximation with heavier tails I: non-asymptotic bounds with weak assumptions and stochastic constraints

We derive new and improved non-asymptotic deviation inequalities for the sample average approximation (SAA) of an optimization problem. Our results give strong error probability bounds that are "sub-Gaussian"~even when the randomness of the problem is fairly heavy tailed. Additionally, we obtain good (often optimal) dependence on the sample size and geometrical parameters of the problem. Finally, we allow for random constraints on the SAA and unbounded feasible sets, which also do not seem to have been considered before in the non-asymptotic literature. Our proofs combine different ideas of potential independent interest: an adaptation of Talagrand's "generic chaining"~bound for sub-Gaussian processes; "localization"~ideas from the Statistical Learning literature; and the use of standard conditions in Optimization (metric regularity, Slater-type conditions) to control fluctuations of the feasible set.

math.OC

Extragradient method with variance reduction for stochastic variational inequalities

We propose an extragradient method with stepsizes bounded away from zero for stochastic variational inequalities requiring only pseudo-monotonicity. We provide convergence and complexity analysis, allowing for an unbounded feasible set, unbounded operator, non-uniform variance of the oracle and, also, we do not require any regularization. Alongside the stochastic approximation procedure, we iteratively reduce the variance of the stochastic error. Our method attains the optimal oracle complexity $\mathcal{O}(1/\epsilon^2)$ (up to a logarithmic term) and a faster rate $\mathcal{O}(1/K)$ in terms of the mean (quadratic) natural residual and the D-gap function, where $K$ is the number of iterations required for a given tolerance $\epsilon>0$. Such convergence rate represents an acceleration with respect to the stochastic error. The generated sequence also enjoys a new feature: the sequence is bounded in $L^p$ if the stochastic error has finite $p$-moment. Explicit estimates for the convergence rate, the oracle complexity and the $p$-moments are given depending on problem parameters and distance of the initial iterate to the solution set. Moreover, sharper constants are possible if the variance is uniform over the solution set or the feasible set. Our results provide new classes of stochastic variational inequalities for which a convergence rate of $\mathcal{O}(1/K)$ holds in terms of the mean-squared distance to the solution set. Our analysis includes the distributed solution of pseudo-monotone Cartesian variational inequalities under partial coordination of parameters between users of a network.

math.OC

Variance-based stochastic extragradient methods with line search for stochastic variational inequalities

A dynamic sampled stochastic approximated (DS-SA) extragradient method for stochastic variational inequalities (SVI) is proposed that is \emph{robust} with respect to an unknown Lipschitz constant $L$. To the best of our knowledge, it is the first provably convergent \emph{robust} SA \emph{method with variance reduction}, either for SVIs or stochastic optimization, assuming just an unbiased stochastic oracle in a large sample regime. This widens the applicability and improves, up to constants, the desired efficient acceleration of previous variance reduction methods, all of which still assume knowledge of $L$ (and, hence, are not robust against its estimate). Precisely, compared to the iteration and oracle complexities of $\mathcal{O}(\epsilon^{-2})$ of previous robust methods with a small stepsize policy, our robust method obtains the faster iteration complexity of $\mathcal{O}(\epsilon^{-1})$ with oracle complexity of $(\ln L)\mathcal{O}(d\epsilon^{-2})$ (up to logs). This matches, up to constants, the sample complexity of the sample average approximation estimator which does not assume additional problem information (such as $L$). Differently from previous robust methods for ill-conditioned problems, we allow an unbounded feasible set and an oracle with multiplicative noise (MN) whose variance is not necessarily uniformly bounded. These properties are seen in our complexity estimates which depend only on $L$ and local second or forth moments at solutions. The robustness and variance reduction properties of our DS-SA line search scheme come at the expense of nonmartingale-like dependencies (NMD) due to the needed inner statistical estimation of a lower bound for $L$. In order to handle a NMD and a MN, our proofs rely on a novel localization argument based on empirical process theory. We also propose another robust method for SVIs over the wider class of H\"older continuous operators.

math.OC

Incremental constraint projection methods for monotone stochastic variational inequalities

We consider stochastic variational inequalities with monotone operators defined as the expected value of a random operator. We assume the feasible set is the intersection of a large family of convex sets. We propose a method that combines stochastic approximation with incremental constraint projections meaning that at each iteration, a step similar to some variant of a deterministic projection method is taken after the random operator is sampled and a component of the intersection defining the feasible set is chosen at random. Such sequential scheme is well suited for applications involving large data sets, online optimization and distributed learning. First, we assume that the variational inequality is weak-sharp. We provide asymptotic convergence, feasibility rate of $O(1/k)$ in terms of the mean squared distance to the feasible set and solvability rate of $O(1/\sqrt{k})$ (up to first order logarithmic terms) in terms of the mean distance to the solution set for a bounded or unbounded feasible set. Then, we assume just monotonicity of the operator and introduce an explicit iterative Tykhonov regularization to the method. We consider Cartesian variational inequalities so as to encompass the distributed solution of stochastic Nash games or multi-agent optimization problems under a limited coordination. We provide asymptotic convergence, feasibility rate of $O(1/k)$ in terms of the mean squared distance to the feasible set and, in the case of a compact set, we provide a near-optimal solvability convergence rate of $O\left(\frac{k^\delta\ln k}{\sqrt{k}}\right)$ in terms of the mean dual gap-function of the SVI for arbitrarily small $\delta>0$.

math.OC