arXiv ScienceSearch

arXiv · 2505.10624

Targeted Learning Estimation of Sampling Variance for Improved Inference

Abstract

For robust statistical inference it is crucial to obtain a good estimator of the variance of the proposed estimator of the statistical estimand. A commonly used estimator of the variance for an asymptotically linear estimator is the sample variance of the estimated influence function. This estimator has been shown to be anti-conservative in limited samples or in the presence of near-positivity violations, leading to elevated Type-I error rates and poor coverage. In this paper, capitalizing on earlier attempts at targeted variance estimators, we propose a one-step targeted variance estimator for the causal risk ratio (CRR) in scenarios involving treatment, outcome, and baseline covariates. While our primary focus is on the variance of log(CRR), our methodology can be extended to other causal effect parameters. Specifically, we focus on the variance of the IF for the log relative risk (log(CRR)) estimator, which requires deriving the efficient influence function for the variance of the IF as the basis for constructing the estimator. Several methods are available to develop efficient estimators of asymptotically linear parameters. In this paper, we concentrate on the so-called one-step targeted maximum likelihood estimator, which is a substitution estimator that utilizes a one-dimensional universal least favorable parametric submodel when updating the distribution. We conduct simulations with different effect sizes, sample sizes and levels of positivity to compare the estimator with existing methods in terms of coverage and Type-I error. Simulation results demonstrate that, especially with small samples and near-positivity violations, the proposed variance estimator offers improved performance, achieving coverage closer to the nominal level of 0.95 and a lower Type-I error rate.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yunwen Ji, Mark van der Laan, Alan Hubbard. 2025-05-15. Targeted Learning Estimation of Sampling Variance for Improved Inference. https://arxiv.org/abs/2505.10624

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Parsimonious Compactly Supported Covariance Models in the Gauss Hypergeometric Class: Identifiability, Reparameterizations, and Asymptotic Properties

We study covariance functions in the Gauss hypergeometric ($\mathcal{GH}$) class, a flexible family that encompasses the Generalized Wendland ($\mathcal{GW}$) and Matérn ($\mathcal{MT}$) models. We derive sharp validity conditions, providing a complete characterization of the admissible parameter space, and show that the model exhibits structural identifiability issues under both increasing- and fixed-domain asymptotics. To resolve this issue, we introduce a parsimonious compactly supported subclass selected via a maximum integral range criterion. The resulting hypergeometric model can be viewed as a structural refinement of the $\mathcal{GW}$ family and admits compact-support reparameterizations that recover the $\mathcal{MT}$ model as a limit case. We further establish strong consistency and asymptotic normality of the maximum likelihood estimator of the associated microergodic parameter under fixed-domain asymptotics. Simulation experiments and a real-data application to climate data illustrate the finite-sample behavior and practical performance of the proposed model.

stat.ME

Doubly Robust Estimation of Treatment Effects in Staggered Difference-in-Differences with Time-Varying Covariates

The difference-in-differences (DiD) design is a quasi-experimental method for estimating treatment effects from panel data. With staggered adoption, the average treatment effect on the treated (ATT) is defined at the group-period level and aggregated into groupwise, periodwise, dynamic, and overall estimands. Existing doubly robust estimators typically compare treated groups to a single reference group and do not account for differing variability across multiple not-yet-treated cohorts. We propose an augmented inverse variance weighting (AIVW) estimator that pools information across all available not-yet-treated groups using conditional-variance weights, extending this approach to settings with time-varying covariates under an explicit covariate-exogeneity condition. Under a homoskedastic working model, AIVW reduces to an augmented inverse probability weighting (AIPW) estimator that is simpler to compute and more robust in finite samples. Both estimators are doubly robust, and we characterize the specific conditions under which they attain the semiparametric efficiency bound. In simulation studies, AIVW and AIPW achieve lower bias and variance than existing doubly robust estimators when not-yet-treated groups differ in conditional variability. As an illustration, we study the effect of a parallel college admission mechanism, relative to immediate admission, on justified envy using staggered provincial reform data from the China National College Entrance Examination.

stat.ME

Spectral Design of Random-Duration Switchbacks

A switchback experiment alternates an entire system between treatment and control over time. It is especially useful when interactions between units can undermine standard unit-level experiments. Switchback experiments are commonly implemented on fixed temporal grids, which impose highly structured restrictions on when treatment can switch. We study a broader class of random-duration switchbacks, in which treatment and control alternate across runs whose durations are drawn from a common distribution. Under finite carryover, we show that the mean squared error of the Horvitz-Thompson estimator has a simple frequency-domain representation governed by how temporal outcome patterns align with design-induced imbalance and contamination patterns. This representation yields a tractable worst-case design criterion that can be computed directly from the duration distribution. Optimizing over even a simple two-rate family produces a run-age-dependent switching rule that reduces the asymptotic worst-case mean squared error by at least 32\% relative to the standard independently randomized fixed-block switchback.

stat.ME