Valid inference for regression with best subset selection
Best subset selection is widely implemented in statistical software and is routinely used by practitioners in scientific fields for variable selection. However, classical confidence intervals and $p$-values constructed after model selection generally fail to achieve their nominal frequentist guarantees, which can invalidate subsequent findings. In this article, we characterize the altered conditional sampling distributions of pivotal quantities after best subset selection. Building on selective-inference techniques developed in related settings, our finite-sample characterization of the AIC selection event reveals that its geometry is a union of finitely many intervals on the real line. This geometry enables exact conditioning and avoids the excessive conditioning common in other post-selection methods. We use this characterization to develop valid inference procedures that provide $p$-values with nominal Type~I error and confidence intervals with finite-sample coverage guarantees. The proposed methods are easy to implement, computationally efficient, and broadly applicable to other commonly used best subset selection criteria. We also study inference with unknown noise level using a Monte Carlo selective test conditional on the AIC-selected model, which controls finite-sample Type~I error at the nominal level under the selected model null. In an application to a classical U.S. consumption dataset, the proposed confidence intervals lead to different conclusions from the conventional intervals, even when the selected model is the full model, producing interpretable findings that better align with empirical observations.