arXiv Science⌕ Search

arXiv · 2609.32910

The Complexity of Convex Optimization with Mismatched Geometry

Abstract

Optimal first-order methods on non-Euclidean domains such as the $\ell_1$ ball $B_1^n(R)=\{x\in\mathbb R^n:\|x\|_1\le R\}$ pair the prox-function with the norm in which smoothness is measured. When the gradient is $L$-Lipschitz in the Euclidean norm only, the accelerated method with a Euclidean prox-setup reduces the functional gap $f(x_N)-\min_{B_1^n(R)}f$ to $O(LR^2/N^2)$ after $N$ first-order queries (an entropic $\ell_1$ prox-setup replaces $L$ by the $\ell_1\to\ell_\infty$ constant $L_1\le L$, at the cost of a factor $\log n$ and with the same exponent). The lower bound of Guzmán and Nemirovski is of order $LR^2/N^3$, and whether the upper bound can be brought down to that order is a question of A. S. Nemirovski. We show that for this mismatched problem the minimax value of the gap is of order $LR^2/N^3$ up to logarithmic factors, for deterministic and for randomized methods alike, and already in dimension proportional to the number of queries. The upper bound is attained by a Steiner-point level method: every supporting hyperplane seen so far is kept as a level cut, and the next query is made near the Steiner point of the resulting localization polytope. The analysis rests on a single geometric fact: along nested subsets of the ball the Steiner points travel a distance that is polylogarithmic in the dimension, in contrast to $\sqrt n$ for the Euclidean ball. All queries stay feasible, and a randomized selector keeps the internal work polynomial. The same geometry yields optimal rates for nonsmooth objectives, Hölder gradients, higher-order oracles and Lipschitz monotone operators, the last with a matching deterministic lower bound. For convex quadratics, a curvature-learning method attains the optimal rate on every $\ell_p$ ball with $1\le p<2$ and without logarithmic loss. Experiments confirm the predicted $N^{-3}$ behaviour on objectives that are hard for Euclidean methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Timofei Loginov, Alexander Gasnikov, Yuriy Dorn, Aleksandr Shestakov, Nazarii Tupitsa, Osman Osmanov, Darina Dvinskikh. 2026-09-26. The Complexity of Convex Optimization with Mismatched Geometry. https://arxiv.org/abs/2609.32910

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Verifiable constraint qualifications for infinite-dimensional optimization problems and their applications to optimal control problems

This paper employs a finite codimensionality condition to establish an enhanced Fritz John condition for general constrained nonlinear infinite-dimensional optimization problems. In applications, this approach provides a new, unified framework for deriving first-order necessary conditions across a broad class of optimal control problems involving both deterministic and stochastic systems. Compared to existing constraint qualifications, our finite codimensionality condition, which is equivalent to the validity of certain \textit{a priori} estimates, yields a direct and analytically tractable verification method for each application. Furthermore, the core ideas of this method can be further extended to investigate the KKT conditions for infinite-dimensional optimization problems.

math.OC↗

Arrival-Intensity Control for A Single-Server Queue in Heavy Traffic

We study a single-server queue control problem (QCP) in a heavy-traffic regime, extending the framework from Lee and Weerasinghe (2011). The state process represents the offered waiting time. Service times and patience times form independent i.i.d. sequences with general distributions. We formulate an infinite-horizon discounted QCP that balances the cost of controlling the arrival intensity against a penalty for server idleness. A distinctive feature of the formulation is the nonstandard decreasing operational running cost arising from the arrival-intensity control mechanism. Under suitable heavy-traffic assumptions, the diffusion-scaled offered waiting-time process converges to a regulated diffusion, leading to an associated diffusion control problem (DCP). We find the optimal control of the associated DCP by incorporating the Legendre-Fenchel transform and a formal Hamilton-Jacobi-Bellman (HJB) equation. We then construct a sequence of intensity controls for the prelimit queueing systems from the DCP-optimal feedback and establish its asymptotic optimality within the specified heavy-traffic admissible control class. Beyond theoretical analysis, numerical experiments further examine whether reinforcement learning can approximate the optimal policy with discounted costs close to the HJB reference. Policies are trained from simulated state transitions and realized costs, and are evaluated against the independently computed HJB feedback.

math.OC↗

Information-Theoretic Upper Bounds for Deterministic Noise in Zeroth-Order Convex Optimization

We study zeroth-order convex optimization on Euclidean balls with function values corrupted by a fixed, uniformly bounded deterministic perturbation. For a query budget $T$, the maximum admissible level of noise (MALN) is the supremum of noise levels for which an $\eps$-accurate solution can be found with probability $1-β$. We construct an explicit hard family: the support function of a spherical belt combined with an affine branch along a hidden random direction. It yields finite-budget upper bounds on the MALN for Lipschitz, strongly convex, uniformly convex, smooth, Hölder, and smooth strongly convex classes. For Lipschitz objectives the bound has order $\eps^2\sqrt\ell/(\sqrt nMR)+\eps\ell/n$, where $\ell=\log(2(T+1)/(1-β))$ and $n-1\ge32\ell$; this is the scale of Li and Risteski with explicit logarithmic factors. A tangential-gradient estimate for the Moreau envelope yields the smooth scale $\eps^{3/2}/(\sqrt n\sqrt L\,R)$ and, for $L-μ\ge5μ$, the smooth strongly convex scale $\eps\sqrt{μ/((L-μ)n)}$. These bounds are tight: on explicit parameter ranges and subclasses, polynomial-query two-point methods tolerate noise of the order of the terms of order $n^{-1/2}$, so that the MALN is determined up to $O(\sqrt\ell)$ and absolute constants whenever these terms dominate. At $L=μ$, simplex interpolation with $n+1$ queries tolerates noise $R\sqrt{2μ\eps/n}$, which is optimal up to $O(\sqrt\ell)$ for $\eps\leμR^2/3$, while with at most $n$ exact queries no randomized algorithm guarantees success probability greater than $1/2$ uniformly over the class. We also quantify noise tolerance near this endpoint and illustrate the hiding mechanism and the breakdown of a two-point method numerically.

math.OC↗