arXiv · 2609.26213
Improved Multiplayer Bandit Algorithm for Bernoulli Rewards
Abstract
We study the multiplayer multi-armed bandit problem with information asymmetry under Bernoulli rewards, for three information structures: asymmetry in actions, in rewards, and in both. Replacing the Hoeffding-style confidence intervals of prior work with Kullback--Leibler (KL) divergence-based bounds gives strictly tighter regret guarantees in each case. We propose \texttt{mKL-UCB}, \texttt{mKL-UCB-Intervals} and \texttt{mKL-DSEE}, and show that the improvement factor is at least two by Pinsker's inequality and far larger when reward means are near zero or one. For asymmetry in rewards we prove that two arms' KL intervals separate after a deterministic number of samples, and that $M$ independent players accelerate elimination further.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Khang Nguyen, Ricardo Parada, William Chang. 2026-08-11. Improved Multiplayer Bandit Algorithm for Bernoulli Rewards. https://arxiv.org/abs/2609.26213
Cite the original work for its findings. Save a collection to share your selection of sources.