arXiv · 2307.00863
Thompson Sampling under Bernoulli Rewards with Local Differential Privacy
Abstract
This paper investigates the problem of regret minimization for multi-armed bandit (MAB) problems with local differential privacy (LDP) guarantee. Given a fixed privacy budget $\epsilon$, we consider three privatizing mechanisms under Bernoulli scenario: linear, quadratic and exponential mechanisms. Under each mechanism, we derive stochastic regret bound for Thompson Sampling algorithm. Finally, we simulate to illustrate the convergence of different mechanisms under different privacy budgets.
Explore related subjects
Keep this discovery
Bo Jiang, Tianchi Zhao, Ming Li. 2023-07-03. Thompson Sampling under Bernoulli Rewards with Local Differential Privacy. https://arxiv.org/abs/2307.00863
Cite the original work for its findings. Save a collection to share your selection of sources.