Nonparametric Bayesian Policy Learning
I propose Nonparametric Bayesian Policy Learning (NBPL) as a framework for uncertainty-aware treatment choice. The key observation is that welfare is fully determined by a reduced-form distribution, so uncertainty about optimal policies entirely reflects uncertainty about this distribution. NBPL places a Dirichlet process prior on the reduced-form distribution and uses the resulting posterior for both traditional policy choice and inference on optimal welfare and optimal treatment assignments. NBPL is computationally tractable: the default Bayesian-bootstrap implementation requires only exponential reweighting of the observations. I establish two theoretical properties. First, posterior welfare regret converges at the minimax-optimal rate, providing a novel policy-relevant analogue of posterior contraction rates. Second, posterior model selection across policy classes is consistent. I relate NBPL to existing policy learning approaches and illustrate using two empirical applications: the JTPA experiment and the bednet subsidy experiment. In both applications, decision-tree rules tend to yield higher optimal welfare than linear rules.