arXiv Science⌕ Search

arXiv · 2610.01572

Optimal Momentum Methods for Stochastic Multilevel Compositional Optimization

Abstract

This paper investigates stochastic multi-level optimization where the objective is a nested composition of several smooth non-convex functions. We assume that only stochastic estimates of the gradient and function values for each level are accessible. Consequently, obtaining an accurate estimate of the overall gradient is challenging due to the nested structure. To address this, we employ a momentum-based estimator with mini-batches to track the function values of each level, which are subsequently used to construct momentum gradient estimators. We establish an optimal sample complexity of $\mathcal{O}(ε^{-4})$ for finding an $ε$-stationary point, avoiding the stronger average smoothness assumption commonly relied upon in prior literature. Furthermore, by employing a normalization technique, we attain the same rate without requiring problem-dependent constants to set hyperparameters. To achieve the optimal rate without mini-batches, we further develop a batch-free method that incorporates a first-order approximation and a clipping technique for function value estimation. Finally, we validate the effectiveness of our proposed methods through experiments on risk-averse portfolio optimization and hierarchical tilted empirical risk minimization.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wei Jiang, Rui Yan, Sifan Yang, Yuanyu Wan, Lijun Zhang, Zechao Li. 2026-10-01. Optimal Momentum Methods for Stochastic Multilevel Compositional Optimization. https://arxiv.org/abs/2610.01572

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Using Less for More: When Warm-Starting Accelerates Branch-and-Cut for Stochastic Programs

Two-stage stochastic programs quickly become intractable as the number of scenarios grows. Motivated by this, we propose TULIP, a modular and easy-to-implement three-step warm-start framework for two-stage stochastic (mixed-)integer programs with an exponential number of cuts separated during branch-and-cut. TULIP (a) builds a cheap surrogate of the full problem by reducing the scenario set or by decoupling the two stages, (b) solves it up to a first incumbent to collect the tight cuts separated along the way, and (c) injects them to warm-start the original problem. In short: we use less (a cheaper surrogate) for more (the original problem). Using this modular setup, we propose four methods within this framework, each with a slightly different setting. Across four case studies, we show that this acceleration is governed by a single mechanism, the root cut loop, and we specify it through a closed-form equation. This TULIP speedup model predicts a speedup when the time saved in the root cut loop exceeds the surrogate overhead. In our experiments, a TULIP variant achieves mean speedups of up to 2.56, with gains increasing with the scenario count. In the remaining case studies, TULIP provides little or no runtime benefit, which the TULIP speedup model mostly explains through insufficient root cut loop savings compared to the surrogate overhead.

math.OC↗

Convergence Analysis of the Wasserstein Proximal Algorithm beyond Geodesic Convexity

The proximal algorithm is a powerful tool to minimize nonlinear and nonsmooth functionals in a general metric space. Motivated by the recent progress in studying the training dynamics of the noisy gradient descent algorithm on two-layer neural networks in the mean-field regime, we provide in this paper a simple and self-contained analysis for the convergence of the general-purpose Wasserstein proximal algorithm without assuming geodesic convexity of the objective functional. Under a natural Wasserstein analog of the Euclidean Polyak-Łojasiewicz inequality, we establish that the proximal algorithm achieves an unbiased and linear convergence rate. Our convergence rate improves upon existing rates of the proximal algorithm for solving Wasserstein gradient flows under strong geodesic convexity. We also extend our analysis to the inexact proximal algorithm for geodesically semiconvex objectives. In our numerical experiments, proximal training demonstrates a faster convergence rate than the noisy gradient descent algorithm on mean-field neural networks.

math.OC↗

Technological foundations of management decision-making in the reconstruction of complex gas pipeline system

This monograph presents a comprehensive analysis of the technological foundations of management decision-making in the reconstruction of complex gas pipeline systems. The study addresses the challenges posed by the aging infrastructure of gas supply networks and explores advanced strategies to improve their reliability, efficiency, and automation. Particular attention is given to the reconstruction of pipelines with various configurations linear, looped, and parallel systems under non-stationary gas flow conditions. The proposed models and methodologies offer solutions for optimizing operational parameters, improving emergency valve response, and ensuring uninterrupted gas supply through advanced management systems and data-driven decision support tools. Emphasis is placed on the integration of modern technologies, system theory, and feedback mechanisms in the design and operation of reconstructed pipeline systems. This work is intended for engineers, system designers, and researchers in the fields of gas supply, systems engineering, and energy infrastructure.

math.OC↗