arXiv · 2609.37380
Information Limits of Multistage Inventory Control: Learning, Valuation, and Censoring
Abstract
Learning how to replenish inventory and estimating the cost of doing so require different information. We establish information-theoretic lower bounds for offline multistage inventory control with independent bounded demand. For fixed additive error in total expected cost, policy learning requires cubically many scalar observations in the horizon in the worst case, even under full observation and adaptive sample allocation. This matches the known empirical-risk upper order. Stationarity permits a quadratic policy-learning guarantee, yet valuation with inherited stock can remain cubically difficult even when the optimal policy is known. Censoring sharpens this distinction: optimal decisions can be learnable while absolute costs remain unidentified. We give a carrysafe coverage condition under which truncation preserves inventory transitions and policy gaps exactly. Combining this reduction with a matching lower bound yields the sharp censored-learning rate, with raw sample requirements inversely proportional to the usable fraction of sales logs. The same qualified observations can check coverage and fit the policy under a joint error guarantee. The results isolate information limits for decisions and valuation under an explicit conditionally independent logging protocol.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hanzhang Qin, David Simchi-Levi, Ruihao Zhu. 2026-09-29. Information Limits of Multistage Inventory Control: Learning, Valuation, and Censoring. https://arxiv.org/abs/2609.37380
Cite the original work for its findings. Save a collection to share your selection of sources.