arXiv · 2607.29375
The Greedy Advantage in Finite-Horizon Bandits
Abstract
Organizations increasingly rely on sequential experimentation to improve decision-making. While the multi-armed bandit literature has developed algorithms with strong asymptotic regret guarantees, many practical applications operate over finite and externally imposed horizons. Motivated by the finite-horizon setting, we develop a class of regularized greedy algorithms for multi-armed Bernoulli bandits. We derive the first finite-horizon regret envelopes for regularized greedy bandits, showing that finite-horizon regret decomposes into transient exploration costs and a suboptimal convergence term that decays exponentially with the regularization strength. This characterization yields principled calibration rules for the regularization parameters and, as a limiting case, sharper regret guarantees for the classical greedy policy. Across extensive numerical experiments, calibrated regularized greedy policies consistently match or outperform state-of-the-art algorithms. These results suggest that regularized greedy policies can provide an effective approach for finite-horizon bandit problems.
Explore related subjects
Keep this discovery
Kai Zhou, Michael Lingzhi Li, Kai Wang. 2026-07-31. The Greedy Advantage in Finite-Horizon Bandits. https://arxiv.org/abs/2607.29375
Cite the original work for its findings. Save a collection to share your selection of sources.