arXiv Science⌕ Search

arXiv subjects

Alexander V. Gasnikov

Publications and source records attributed to Alexander V. Gasnikov.

3 recordsLinked to original sources

Saddle-Point Problems with a Low-Dimensional Block Do Not Need Accurate Inner Solves

Many learning problems couple a high-dimensional block of parameters with a handful of adversarial or dual variables: worst-group risk over a few groups, or learning under a few constraints. We consider $\min_{x\in X}\max_{y\in Y} f(x,y)$ with $X\subseteq\mathbb{R}^n$, a compact convex set and strongly convex with condition number $κ$, and we count the oracle calls for $x$ and for $y$ separately. The textbook approach runs a cutting-plane method in $y$ and solves every inner problem to accuracy $\varepsilon$ with an accelerated method; it needs $O(m\log(1/\varepsilon))$ calls for $y$ but $O(m\sqrtκ\log^2(1/\varepsilon))$ calls for $x$. We show that the inner problems need not be solved accurately. Our method, certificate transport, keeps a strongly convex lower model of a single slice and uses it as a prior for a short accelerated run on the next slice. By concavity in $y$, every call for $y$ then either cuts the localizer or moves the lower model to a mixture of the two slices with a certified increase of the lower bound. For $m=1$ this gives an $\varepsilon$-saddle point after $O(\sqrtκ\log(1/\varepsilon))$ calls for $x$, up to an initialization term, and $O(\log(1/\varepsilon))$ calls for $y$; both counts are optimal, even though $f(x,\cdot)$ is only assumed to be concave and Lipschitz. For general $m$, with centers of gravity of the localizer treated as computable, the method needs $O((m+\sqrt{mκ})\log(1/\varepsilon))$ calls for $x$ and the optimal $O(m\log(1/\varepsilon))$ calls for $y$. Under strong concavity in $y$, a two-point accelerated method makes the number of calls for $x$ independent of $m$.

math.OC↗

Application of Optimal Inexact Second-Order Acceleration to Distributed Stochastic Optimization under Statistical Similarity

We consider distributed stochastic convex optimization with a fixed budget of $N$ independent samples split among $m$ workers. Sample average approximation reduces the problem to a regularized finite-sum problem whose local Hessians are statistically similar. This allows the Hessian of the local objective at the server to be used as an inexact Hessian of the global objective, while the workers communicate only gradients. We apply the optimal accelerated inexact Newton extragradient method of (Chen et al., 2026) and propose its distributed restarted variant for the strongly convex empirical problem. The method reaches the statistical accuracy of order $N^{-1/2}$ in $\widetilde O\left(\max\{N^{1/7},m^{1/4}\}\right)$ communication rounds. Hence, with $m=N^{4/7}$ workers, it requires $\widetilde O\left(N^{1/7}\right)$ rounds, improving the dependence on the total sample size from $\widetilde O\left(N^{1/6}\right)$ for the previous accelerated cubic Newton construction of (Agafonov et al., 2021). Each iteration uses two gradient aggregation rounds and does not require Hessian communication.

math.OC↗

Certified Residual Quasi-Newton Methods for Distributed Variational Inequalities

Second-order methods for smooth monotone variational inequalities reach the optimal rate $O(T^{-3/2})$, but a distributed exact Jacobian costs $d$ times more communication than an operator value. We show that similarity does part of the work for free: if the server's Jacobian differs from the global one by at most $β$, using it gives $O(L_1D^3T^{-3/2}+βD^2T^{-1})$ at first-order communication cost. A quasi-Newton approximation of the residual Jacobian $\nabla F-\nabla F_1$, built from secants already communicated, improves the model but cannot remove the $T^{-1}$ term, because any uniform bound on the Jacobian error leaves it in the rate. We therefore certify the surrogate only along the candidate step: one Jacobian-vector product tests it, and a failed test is reused as an exact correction. This attains the exact rate $O(L_1D^3T^{-3/2})$ while transmitting only vectors. Experiments on LIBSVM and synthetic instances measure accuracy against communication.

math.OC↗