arXiv ScienceSearch

arXiv subjects

Arkadi Nemirovski

Publications and source records attributed to Arkadi Nemirovski.

At least 37 records · Page 2Linked to original sources

On Polyhedral Estimation of Signals via Indirect Observations

We consider the problem of recovering linear image of unknown signal belonging to a given convex compact signal set from noisy observation of another linear image of the signal. We develop a simple generic efficiently computable nonlinear in observations "polyhedral" estimate along with computation-friendly techniques for its design and risk analysis. We demonstrate that under favorable circumstances the resulting estimate is provably near-optimal in the minimax sense, the "favorable circumstances" being less restrictive than the weakest known so far assumptions ensuring near-optimality of estimates which are linear in observations.

math.ST

Hypothesis Testing via Euclidean Separation

We discuss an "operational" approach to testing convex composite hypotheses when the underlying distributions are heavy-tailed. It relies upon Euclidean separation of convex sets and can be seen as an extension of the approach to testing by convex optimization developed in [8, 12]. In particular, we show how one can construct quasi-optimal testing procedures for families of distributions which are majorated, in a certain precise sense, by a sub-spherical symmetric one and study the relationship between tests based on Euclidean separation and "potential-based tests." We apply the promoted methodology in the problem of sequential detection and illustrate its practical implementation in an application to sequential detection of changes in the input of a dynamic system. [8] Goldenshluger, Alexander and Juditsky, Anatoli and Nemirovski, Arkadi, Hypothesis testing by convex optimization, Electronic Journal of Statistics,9 (2):1645-1712, 2015. [12] Juditsky, Anatoli and Nemirovski, Arkadi, Hypothesis testing via affine detectors, Electronic Journal of Statistics, 10:2204--2242, 2016.

math.ST

Near-Optimality of Linear Recovery from Indirect Observations

We consider the problem of recovering linear image $Bx$ of a signal $x$ known to belong to a given convex compact set ${\cal X}$ from indirect observation $\omega=Ax+\xi$ of $x$ corrupted by random noise $\xi$ with finite covariance matrix. It is shown that under some assumptions on ${\cal X}$ (satisfied, e.g., when ${\cal X}$ is the intersection of $K$ concentric ellipsoids/elliptic cylinders, or the unit ball of the spectral norm in the space of matrices) and on the norm $\|\cdot\|$ used to measure the recovery error (satisfied, e.g., by $\|\cdot\|_p$-norms, $1\leq p\leq 2$, on ${\mathbf{R}}^m$ and by the nuclear norm on the space of matrices), one can build, in a computationally efficient manner, a "presumably good" linear in observations estimate, and that in the case of zero mean Gaussian observation noise, this estimate is near-optimal among all (linear and nonlinear) estimates in terms of its worst-case, over $x\in {\cal X}$, expected $\|\cdot\|$-loss. These results form an essential extension of those in our paper arXiv:1602.01355, where the assumptions on ${\cal X}$ were more restrictive, and the norm $\|\cdot\|$ was assumed to be the Euclidean one. In addition, we develop near-optimal estimates for the case of "uncertain-but-bounded" noise, where all we know about $\xi$ is that it is bounded in a given norm by a given $\sigma$. Same as in arXiv:1602.01355, our results impose no restrictions on $A$ and $B$. This arXiv paper slightly strengthens the journal publication Juditsky, A., Nemirovski, A. "Near-Optimality of Linear Recovery from Indirect Observations," Mathematical Statistics and Learning 1:2 (2018), 171-225.

math.ST

Estimating Linear and Quadratic forms via Indirect Observations

In this paper, we further develop the approach, originating in [14 (arXiv:1311.6765),20 (arXiv:1604.02576)], to "computation-friendly" hypothesis testing and statistical estimation via Convex Programming. Specifically, we focus on estimating a linear or quadratic form of an unknown "signal," known to belong to a given convex compact set, via noisy indirect observations of the signal. Most of the existing theoretical results on the subject deal with precisely stated statistical models and aim at designing statistical inferences and quantifying their performance in a closed analytic form. In contrast to this descriptive (and highly instructive) traditional framework, the approach we promote here can be qualified as operational -- the estimation routines and their risks are yielded by an efficient computation. All we know in advance is that under favorable circumstances to be specified below, the risk of the resulting estimate, whether high or low, is provably near-optimal under the circumstances. As a compensation for the lack of "explanatory power," this approach is applicable to a much wider family of observation schemes than those where "closed form descriptive analysis" is possible. The paper is a follow-up to our paper [20 (arXiv:1604.02576)] dealing with hypothesis testing, in what follows, we apply the machinery developed in this reference to estimating linear and quadratic forms.

math.ST

Change Detection via Affine and Quadratic Detectors

The goal of the paper is to develop a specific application of the convex optimization based hypothesis testing techniques developed in A. Juditsky, A. Nemirovski, "Hypothesis testing via affine detectors," Electronic Journal of Statistics 10:2204--2242, 2016. Namely, we consider the Change Detection problem as follows: given an evolving in time noisy observations of outputs of a discrete-time linear dynamical system, we intend to decide, in a sequential fashion, on the null hypothesis stating that the input to the system is a nuisance, vs. the alternative stating that the input is a "nontrivial signal," with both the nuisances and the nontrivial signals modeled as inputs belonging to finite unions of some given convex sets. Assuming the observation noises zero mean sub-Gaussian, we develop "computation-friendly" sequential decision rules and demonstrate that in our context these rules are provably near-optimal.

math.ST

Structure-Blind Signal Recovery

We consider the problem of recovering a signal observed in Gaussian noise. If the set of signals is convex and compact, and can be specified beforehand, one can use classical linear estimators that achieve a risk within a constant factor of the minimax risk. However, when the set is unspecified, designing an estimator that is blind to the hidden structure of the signal remains a challenging problem. We propose a new family of estimators to recover signals observed in Gaussian noise. Instead of specifying the set where the signal lives, we assume the existence of a well-performing linear estimator. Proposed estimators enjoy exact oracle inequalities and can be efficiently computed through convex optimization. We present several numerical illustrations that show the potential of the approach.

math.ST

Hypothesis Testing via Affine Detectors

In this paper, we further develop the approach, originating in [GJN], to "computation-friendly" hypothesis testing via Convex Programming. Most of the existing results on hypothesis testing aim to quantify in a closed analytic form separation between sets of distributions allowing for reliable decision in precisely stated observation models. In contrast to this descriptive (and highly instructive) traditional framework, the approach we promote here can be qualified as operational -- the testing routines and their risks are yielded by an efficient computation. All we know in advance is that, under favorable circumstances, specified in [GJN], the risk of such test, whether high or low, is provably near-optimal under the circumstances. As a compensation for the lack of "explanatory power," this approach is applicable to a much wider family of observation schemes and hypotheses to be tested than those where "closed form descriptive analysis" is possible. In the present paper our primary emphasis is on computation: we make a step further in extending the principal tool developed in the cited paper -- testing routines based on affine detectors -- to a large variety of testing problems. The price of this development is the loss of blanket near-optimality of the proposed procedures (though it is still preserved in the observation schemes studied in [GJN], which now become particular cases of the general setting considered here). [GJN]: Goldenshluger, A., Juditsky, A., Nemirovski, A. "Hypothesis testing by convex optimization," Electronic Journal of Statistics 9(2), 2015

math.ST

Near-Optimality of Linear Recovery in Gaussian Observation Scheme under $\|\cdot\|_2^2$-Loss

We consider the problem of recovering linear image $Bx$ of a signal $x$ known to belong to a given convex compact set $X$ from indirect observation $\omega=Ax+\sigma\xi$ of $x$ corrupted by Gaussian noise $\xi$. It is shown that under some assumptions on $X$ (satisfied, e.g., when $X$ is the intersection of $K$ concentric ellipsoids/elliptic cylinders), an easy-to-compute linear estimate is near-optimal, in certain precise sense, in terms of its worst-case, over $x\in X$, expected $\|\cdot\|_2^2$-error. The main novelty here is that our results impose no restrictions on $A$ and $B$, to the best of our knowledge, preceding results on optimality of linear estimates dealt either with the case of direct observations $A=I$ and $B=I$, or with the "diagonal case" where $A$, $B$ are diagonal and $X$ is given by a "separable" constraint like $X=\{x:\sum_ia_i^2x_i^2\leq 1\}$ or $X=\{x:\max_i|a_ix_i|\leq1\}$, or with estimating a linear form (i.e., the case one-dimensional $Bx$).

math.ST

Non-asymptotic confidence bounds for the optimal value of a stochastic program

We discuss a general approach to building non-asymptotic confidence bounds for stochastic optimization problems. Our principal contribution is the observation that a Sample Average Approximation of a problem supplies upper and lower bounds for the optimal value of the problem which are essentially better than the quality of the corresponding optimal solutions. At the same time, such bounds are more reliable than "standard" confidence bounds obtained through the asymptotic approach. We also discuss bounding the optimal value of MinMax Stochastic Optimization and stochastically constrained problems. We conclude with a simulation study illustrating the numerical behavior of the proposed bounds.

math.OC

Decomposition Techniques for Bilinear Saddle Point Problems and Variational Inequalities with Affine Monotone Operators on Domains Given by Linear Minimization Oracles

The majority of First Order methods for large-scale convex-concave saddle point problems and variational inequalities with monotone operators are proximal algorithms which at every iteration need to minimize over problem's domain X the sum of a linear form and a strongly convex function. To make such an algorithm practical, X should be proximal-friendly -- admit a strongly convex function with easy to minimize linear perturbations. As a byproduct, X admits a computationally cheap Linear Minimization Oracle (LMO) capable to minimize over X linear forms. There are, however, important situations where a cheap LMO indeed is available, but X is not proximal-friendly, which motivates search for algorithms based solely on LMO's. For smooth convex minimization, there exists a classical LMO-based algorithm -- Conditional Gradient. In contrast, known to us LMO-based techniques for other problems with convex structure (nonsmooth convex minimization, convex-concave saddle point problems, even as simple as bilinear ones, and variational inequalities with monotone operators, even as simple as affine) are quite recent and utilize common approach based on Fenchel-type representations of the associated objectives/vector fields. The goal of this paper is to develop an alternative (and seemingly much simpler) LMO-based decomposition techniques for bilinear saddle point problems and for variational inequalities with affine monotone operators.

math.OC

On sequential hypotheses testing via convex optimization

We propose a new approach to sequential testing which is an adaptive (on-line) extension of the (off-line) framework developed in [10]. It relies upon testing of pairs of hypotheses in the case where each hypothesis states that the vector of parameters underlying the dis- tribution of observations belongs to a convex set. The nearly optimal under appropriate conditions test is yielded by a solution to an efficiently solvable convex optimization prob- lem. The proposed methodology can be seen as a computationally friendly reformulation of the classical sequential testing.

math.ST

Solving Variational Inequalities with Monotone Operators on Domains Given by Linear Minimization Oracles

The standard algorithms for solving large-scale convex-concave saddle point problems, or, more generally, variational inequalities with monotone operators, are proximal type algorithms which at every iteration need to compute a prox-mapping, that is, to minimize over problem's domain $X$ the sum of a linear form and the specific convex distance-generating function underlying the algorithms in question. Relative computational simplicity of prox-mappings, which is the standard requirement when implementing proximal algorithms, clearly implies the possibility to equip $X$ with a relatively computationally cheap Linear Minimization Oracle (LMO) able to minimize over $X$ linear forms. There are, however, important situations where a cheap LMO indeed is available, but where no proximal setup with easy-to-compute prox-mappings is known. This fact motivates our goal in this paper, which is to develop techniques for solving variational inequalities with monotone operators on domains given by Linear Minimization Oracles. The techniques we develope can be viewed as a substantial extension of the proposed in [5] method of nonsmooth convex minimization over an LMO-represented domain.

math.OC

Mirror Prox Algorithm for Multi-Term Composite Minimization and Semi-Separable Problems

In the paper, we develop a composite version of Mirror Prox algorithm for solving convex-concave saddle point problems and monotone variational inequalities of special structure, allowing to cover saddle point/variational analogies of what is usually called "composite minimization" (minimizing a sum of an easy-to-handle nonsmooth and a general-type smooth convex functions "as if" there were no nonsmooth component at all). We demonstrate that the composite Mirror Prox inherits the favourable (and unimprovable already in the large-scale bilinear saddle point case) $O(1/\epsilon)$ efficiency estimate of its prototype. We demonstrate that the proposed approach can be naturally applied to Lasso-type problems with several penalizing terms (e.g. acting together $\ell_1$ and nuclear norm regularization) and to problems of the structure considered in the alternating directions methods, implying in both cases methods with the $O(\epsilon^{-1})$ complexity bounds.

math.OC

On Lower Complexity Bounds for Large-Scale Smooth Convex Optimization

We derive lower bounds on the black-box oracle complexity of large-scale smooth convex minimization problems, with emphasis on minimizing smooth (with Holder continuous, with a given exponent and constant, gradient) convex functions over high-dimensional ||.||_p-balls, 1<=p<=\infty. Our bounds turn out to be tight (up to logarithmic in the design dimension factors), and can be viewed as a substantial extension of the existing lower complexity bounds for large-scale convex minimization covering the nonsmooth case and the 'Euclidean' smooth case (minimization of convex functions with Lipschitz continuous gradients over Euclidean balls). As a byproduct of our results, we demonstrate that the classical Conditional Gradient algorithm is near-optimal, in the sense of Information-Based Complexity Theory, when minimizing smooth convex functions over high-dimensional ||.||_\infty-balls and their matrix analogies -- spectral norm balls in the spaces of square matrices.

math.OC

Conditional Gradient Algorithms for Norm-Regularized Smooth Convex Optimization

Motivated by some applications in signal processing and machine learning, we consider two convex optimization problems where, given a cone $K$, a norm $\|\cdot\|$ and a smooth convex function $f$, we want either 1) to minimize the norm over the intersection of the cone and a level set of $f$, or 2) to minimize over the cone the sum of $f$ and a multiple of the norm. We focus on the case where (a) the dimension of the problem is too large to allow for interior point algorithms, (b) $\|\cdot\|$ is "too complicated" to allow for computationally cheap Bregman projections required in the first-order proximal gradient algorithms. On the other hand, we assume that {it is relatively easy to minimize linear forms over the intersection of $K$ and the unit $\|\cdot\|$-ball}. Motivating examples are given by the nuclear norm with $K$ being the entire space of matrices, or the positive semidefinite cone in the space of symmetric matrices, and the Total Variation norm on the space of 2D images. We discuss versions of the Conditional Gradient algorithm capable to handle our problems of interest, provide the related theoretical efficiency estimates and outline some applications.

math.OC

Dual subgradient algorithms for large-scale nonsmooth learning problems

"Classical" First Order (FO) algorithms of convex optimization, such as Mirror Descent algorithm or Nesterov's optimal algorithm of smooth convex optimization, are well known to have optimal (theoretical) complexity estimates which do not depend on the problem dimension. However, to attain the optimality, the domain of the problem should admit a "good proximal setup". The latter essentially means that 1) the problem domain should satisfy certain geometric conditions of "favorable geometry", and 2) the practical use of these methods is conditioned by our ability to compute at a moderate cost {\em proximal transformation} at each iteration. More often than not these two conditions are satisfied in optimization problems arising in computational learning, what explains why proximal type FO methods recently became methods of choice when solving various learning problems. Yet, they meet their limits in several important problems such as multi-task learning with large number of tasks, where the problem domain does not exhibit favorable geometry, and learning and matrix completion problems with nuclear norm constraint, when the numerical cost of computing proximal transformation becomes prohibitive in large-scale problems. We propose a novel approach to solving nonsmooth optimization problems arising in learning applications where Fenchel-type representation of the objective function is available. The approach is based on applying FO algorithms to the dual problem and using the {\em accuracy certificates} supplied by the method to recover the primal solution. While suboptimal in terms of accuracy guaranties, the proposed approach does not rely upon "good proximal setup" for the primal problem but requires the problem domain to admit a Linear Optimization oracle -- the ability to efficiently maximize a linear form on the domain of the primal problem.

math.OC

On detecting harmonic oscillations

In this paper, we focus on the following testing problem: assume that we are given observations of a real-valued signal along the grid $0,1,\ldots,N-1$, corrupted by white Gaussian noise. We want to distinguish between two hypotheses: (a) the signal is a nuisance - a linear combination of $d_n$ harmonic oscillations of known frequencies, and (b) signal is the sum of a nuisance and a linear combination of a given number $d_s$ of harmonic oscillations with unknown frequencies, and such that the distance (measured in the uniform norm on the grid) between the signal and the set of nuisances is at least $\rho>0$. We propose a computationally efficient test for distinguishing between (a) and (b) and show that its "resolution" (the smallest value of $\rho$ for which (a) and (b) are distinguished with a given confidence $1-\alpha$) is $\mathrm{O}(\sqrt{\ln(N/\alpha)/N})$, with the hidden factor depending solely on $d_n$ and $d_s$ and independent of the frequencies in question. We show that this resolution, up to a factor which is polynomial in $d_n,d_s$ and logarithmic in $N$, is the best possible under circumstances. We further extend the outlined results to the case of nuisances and signals close to linear combinations of harmonic oscillations, and provide illustrative numerical results.

math.ST

On solving large scale polynomial convex problems by randomized first-order algorithms

One of the most attractive recent approaches to processing well-structured large-scale convex optimization problems is based on smooth convex-concave saddle point reformu-lation of the problem of interest and solving the resulting problem by a fast First Order saddle point method utilizing smoothness of the saddle point cost function. In this paper, we demonstrate that when the saddle point cost function is polynomial, the precise gra-dients of the cost function required by deterministic First Order saddle point algorithms and becoming prohibitively computationally expensive in the extremely large-scale case, can be replaced with incomparably cheaper computationally unbiased random estimates of the gradients. We show that for large-scale problems with favourable geometry, this randomization accelerates, progressively as the sizes of the problem grow, the solution process. This extends significantly previous results on acceleration by randomization, which, to the best of our knowledge, dealt solely with bilinear saddle point problems. We illustrate our theoretical findings by instructive and encouraging numerical experiments.

cs.DS