arXiv · 2609.19677
A Logarithmic Regret Bound for Optimistic Hedge in General-Sum Games
Abstract
Can simple no-regret dynamics attain smaller regret in self-play than against arbitrary adversaries? In $n$-player general-sum games, Daskalakis et al. 2021 proved an $O(n\log d_i\log^4 T)$ individual regret bound for Optimistic Hedge, which improves upon the classical $O(\sqrt T)$ adversarial regret bound. In this work, we show that Optimistic Hedge with a constant step size can further achieve $O(\sqrt n\log d_i\log T)$ individual external regret under expected loss-vector feedback. The time-averaged play consequently enjoys a coarse correlated equilibrium gap $O(\sqrt n\log d\log T/T)$, where $d=\max_i d_i$. The improvement comes from a larger admissible step size $η=Θ(1/(\sqrt n\log T))$. Our analysis proves factorial bounds on high-order differences of probability-weighted pairwise loss gaps, then applies finite-difference interpolation in a fixed Euclidean norm. These estimates sharpen the analysis of Daskalakis et al. 2021 and yield a logarithmic regret bound.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Junsoo Ha. 2026-09-17. A Logarithmic Regret Bound for Optimistic Hedge in General-Sum Games. https://arxiv.org/abs/2609.19677
Cite the original work for its findings. Save a collection to share your selection of sources.