arXiv · 2609.32772
Learnable Randomization as Commitment Against Adaptive Optimizers
Abstract
A pricing page can walk the posted price up to the last amount a buyer still accepts, a recommender can hold back a better item for a barely acceptable promoted one, and a classifier can shift its boundary once applicants change their features. The system predicts the response and then picks the menu that serves its own objective, so the surplus above the user's cutoff is taken. Playing the single best action publishes that cutoff, while noise on actions the user would never take throws away payoff and teaches the platform that a worse menu is still acceptable. We study unpredictable near-optimal policies (UNOP), which mix uniformly on near-best actions that remain individually rational. The mixture is a commitment about the response. On a finite price grid, when the best sure-demand price strictly out-earns the randomized band, a seller who already knows the curve posts below the band, and the purchase that occurs is deterministic. Knowing that curve is not the same as predicting the next draw. The mixture can be learned and the optimizer can match its best response, while the user's payoff stays higher because the mixture changes which action is targeted. In pricing and in policy-aware recommendation this leaves more surplus than greedy play when the platform optimizes against the curve and more than one action is acceptable. The gain goes away under quality ranking, a singleton near-optimal set, a wrong utility estimate, or a short-horizon explorer. That is also where mixing should be turned off if the other side is trying to cooperate.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zihan Deng, Chuanzhi Xu, Xiaozhen Zhong, Haoyang Li, Junjie Huang. 2026-09-26. Learnable Randomization as Commitment Against Adaptive Optimizers. https://arxiv.org/abs/2609.32772
Cite the original work for its findings. Save a collection to share your selection of sources.