Calibrated Estimation and Inference for Semiparametric Regression Models
We consider a broad class of semiparametric regression models in which the conditional distribution of the response takes the form $f\{Y|\boldsymbol{x}^T\boldsymbolβ+m(z),ϕ\}$, known up to a parametric component $\boldsymbolβ$ of diverging dimension $p$, a smooth function $m(\cdot)$, and a dispersion parameter $ϕ$. The existing literature on such models has focused on semiparametric efficiency for $\boldsymbolβ$, treating $ϕ$ and $m(\cdot)$ as nuisances and largely ignoring finite-sample bias. Yet this bias can be substantial, particularly when $p$ is large relative to $n$ or the dispersion is high, and it can seriously undermine inference for $\boldsymbolβ$; moreover, $ϕ$ is often of direct scientific interest. We therefore propose SABRE, a general calibration framework for semiparametric estimation and inference, which calibrates an initial estimator against its model-implied expectation under a tractable parametric approximation to the semiparametric model. For generalized partially linear models, we show that SABRE reduces the bias of both $\boldsymbolβ$ and $ϕ$, accommodates a diverging parameter dimension without sparsity, and preserves the first-order variance and semiparametric efficiency of the initial estimator; the joint construction also improves estimation and inference for $m(\cdot)$. Simulation studies and an application to Alzheimer's disease genetics association analysis demonstrate the empirical effectiveness of SABRE in reducing bias and improving inference.