arXiv ScienceSearch

arXiv subjects

Huiming Lin

Publications and source records attributed to Huiming Lin.

2 recordsLinked to original sources

Valid inference for regression with best subset selection

Best subset selection is widely implemented in statistical software and is routinely used by practitioners in scientific fields for variable selection. However, classical confidence intervals and $p$-values constructed after model selection generally fail to achieve their nominal frequentist guarantees, which can invalidate subsequent findings. In this article, we characterize the altered conditional sampling distributions of pivotal quantities after best subset selection. Building on selective-inference techniques developed in related settings, our finite-sample characterization of the AIC selection event reveals that its geometry is a union of finitely many intervals on the real line. This geometry enables exact conditioning and avoids the excessive conditioning common in other post-selection methods. We use this characterization to develop valid inference procedures that provide $p$-values with nominal Type~I error and confidence intervals with finite-sample coverage guarantees. The proposed methods are easy to implement, computationally efficient, and broadly applicable to other commonly used best subset selection criteria. We also study inference with unknown noise level using a Monte Carlo selective test conditional on the AIC-selected model, which controls finite-sample Type~I error at the nominal level under the selected model null. In an application to a classical U.S. consumption dataset, the proposed confidence intervals lead to different conclusions from the conventional intervals, even when the selected model is the full model, producing interpretable findings that better align with empirical observations.

stat.ME

Double spike Dirichlet priors for structured weighting

Assigning weights to a large pool of objects is a fundamental task in a wide variety of applications. In this article, we introduce the concept of structured high-dimensional probability simplexes, in which most components are zero or near zero and the remaining ones are close to each other. Such structure is well motivated by (i) high-dimensional weights that are common in modern applications, and (ii) ubiquitous examples in which equal weights -- despite their simplicity -- often achieve favorable or even state-of-the-art predictive performance. This particular structure, however, presents unique challenges partly because, unlike high-dimensional linear regression, the parameter space is a simplex and pattern switching between partial constancy and sparsity is unknown. To address these challenges, we propose a new class of double spike Dirichlet priors to shrink a probability simplex to one with the desired structure. When applied to ensemble learning, such priors lead to a Bayesian method for structured high-dimensional ensembles that is useful for forecast combination and improving random forests, while enabling uncertainty quantification. We design efficient Markov chain Monte Carlo algorithms for implementation. Posterior contraction rates are established to study large sample behaviors of the posterior distribution. We demonstrate the wide applicability and competitive performance of the proposed methods through simulations and two real data applications using the European Central Bank Survey of Professional Forecasters data set and a data set from the UC Irvine Machine Learning Repository (UCI).

stat.ME