arXiv · 2609.21262
Optimal Regret for Online Storage Control via Cumulative Policies
Abstract
We study online control of a scalar storage system with adversarial nonnegative arrivals, known retention coefficient, and convex costs depending on both state and action. Each action must respect current resource availability and is chosen before the current arrival and cost function are revealed. For the existing simplex disturbance-action policy class, we give an exact reparameterization by cumulative allocation fractions and a decay-weighted projected subgradient update. The resulting regret bound is independent of policy memory length. For fixed retention coefficient and cost constants, the controller achieves $O(\sqrt T)$ regret against the best fixed infinite-memory policy in this class, using $O(\log T)$ memory and arithmetic operations per round and one cost-subgradient query. A storage-specific block construction gives a matching lower bound against every causal feasible controller, including randomized controllers. Writing $τ=(1-α)^{-1}$, the minimax expected regret is $Θ(\sqrt T\min\{T,τ\}^{3/2})$ for every finite-memory simplex policy class and its infinite-memory extension, when $α\in[1/2,1)$, $T\ge4$, and the positive cost constants are fixed. This identifies the joint horizon and retention-time dependence for these policy benchmarks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kamiar Asgari, Michael J. Neely. 2026-09-18. Optimal Regret for Online Storage Control via Cumulative Policies. https://arxiv.org/abs/2609.21262
Cite the original work for its findings. Save a collection to share your selection of sources.