arXiv ScienceSearch

arXiv subjects

Stijn Vansteelandt

Publications and source records attributed to Stijn Vansteelandt.

At least 19 recordsLinked to original sources

Robust Covariate Adjustment in Multi-Center Randomized Trials

Augmented inverse probability weighting and G-computation with canonical generalized linear models have become increasingly popular for estimating the average treatment effect (ATE) in randomized experiments. These methods leverage outcome prediction models to adjust for imbalances in baseline covariates across treatment arms, improving power compared to unadjusted analyses, while controlling Type I error, even when models are misspecified. In multi-center trials they are often implemented without accounting for clustering by centers. We investigate how ignoring center-level correlation can impair estimation, degrade coverage of confidence intervals, and obscure interpretation. We find these issues to be especially acute for estimators of counterfactual means, as shown through simulations and clarified via theoretical arguments. To address these challenges, we develop semiparametric efficient estimators of counterfactual means and ATE defined for a randomly sampled center and patient. These estimators leverage outcome prediction models to improve efficiency yet retain large-sample unbiasedness under model misspecification. We further introduce an inference framework, inspired by random-effects meta-analysis, tailored to settings with many small centers. Incorporating center effects into the prediction models yields substantial efficiency gains, particularly when treatment effects vary across centers. Simulations and application to the WASH Benefits Bangladesh trial illustrate strong finite-sample performance of the proposed methods.

stat.ME

Automated, efficient and model-free inference for randomized clinical trials via data-driven covariate adjustment

In 2023, the U.S. Food and Drug Administration issued guidance for adjustment of covariates in randomized clinical trials, emphasizing its role in enhancing precision and power through prognostic baseline variables. Despite its potential, many trials underutilize this method partly due to challenges in pre-specifying optimal baseline covariates and their functional forms. We explore the potential of automated, data-adaptive methods-including stepwise regression, Lasso and flexible machine learning algorithms-for covariate adjustment, addressing the challenge of pre-specification. Our approach ensures valid and interpretable treatment effect estimates and standard errors, even when outcome models are misspecified or biased outcome predictions are used. This differs from most competing methods, which assume correctly specified models for consistent standard errors. Our estimators require cross-fitting for reliable standard error estimation, though it can be omitted when variable selection is used, provided the outcome model satisfies an ultra-sparsity assumption. As such, we arrive at simple estimators and standard errors for marginal treatment effects in randomized clinical trials (or similar studies like A/B-testing), exploiting data-adaptive predictions from prognostic baseline covariates, with little (or no) bias in finite samples even when predictions are biased. Empirical and methodological results demonstrate promise of automated covariate adjustment for improving statistical power of trial analyses.

stat.ME

Targeted learning of heterogeneous treatment effect curves for right censored or left truncated time-to-event data

In recent years, there has been growing interest in causal machine learning estimators for quantifying subject-specific effects of a binary treatment on time-to-event outcomes. Estimation approaches have been proposed which attenuate the inherent regularisation bias in machine learning predictions, with each of these estimators addressing measured confounding, right censoring, and in some cases, left truncation. However, the existing approaches are found to exhibit suboptimal finite-sample performance, with none of the existing estimators fully leveraging the temporal structure of the data, yielding non-smooth treatment effects over time. We address these limitations by introducing surv-iTMLE, a targeted learning procedure for estimating the difference in the conditional survival probabilities under two treatments. Unlike existing estimators, surv-iTMLE accommodates both left truncation and right censoring while enforcing smoothness and boundedness of the estimated treatment effect curve over time. Through extensive simulation studies under both right censoring and left truncation scenarios, we demonstrate that surv-iTMLE outperforms existing methods in terms of bias and smoothness of time-varying effect estimates in finite samples. We then illustrate surv-iTMLE's practical utility by exploring heterogeneity in the effects of immunotherapy on survival among non-small cell lung cancer (NSCLC) patients, revealing clinically meaningful temporal patterns that existing estimators may obscure.

stat.ME

Robust evaluation of treatment effects in longitudinal studies with truncation by death or other intercurrent events

Intercurrent events, such as treatment switching, rescue medication, dropout, or truncation by death, frequently complicate intention-to-treat analyses in randomized clinical trials. Existing causal inference frameworks typically target hypothetical or principal stratum estimands (e.g., survivor average causal effects), which rely on unverifiable assumptions and can be sensitive to unmeasured confounders or positivity violations. We propose a novel approach that mitigates this sensitivity by using only information measured prior to the intercurrent event. Our key idea is to compare treated and untreated individuals, matched on baseline covariates, at the most recent time point before either experiences an intercurrent event. We call these contrasts Pairwise Last Observation Time (PLOT) estimands. PLOT estimands are identified in randomized trials without structural assumptions, even under severe positivity violations. Although PLOT-based tests may theoretically be susceptible to residual selection bias, we show this bias vanishes under standard conditions and remains negligible in extensive simulations. We develop asymptotically efficient, model-free tests and treatment effect estimators using data-adaptive nuisance parameter estimation. We evaluate performance via simulation and apply the method to re-analyze the DEVOTE trial, affected by truncation by death. PLOT offers a robust, data-driven alternative for evaluating treatment efficacy in the presence of complex intercurrent events.

stat.ME

Parameterising the effect of a continuous treatment using average derivative effects

The average treatment effect (ATE) is commonly used to quantify the main effect of a binary treatment on an outcome. Extensions to continuous treatments are usually based on the dose-response curve or shift interventions, but both require strong overlap conditions and the resulting curves may be difficult to summarise. We focus instead on average derivative effects (ADEs) that are scalar estimands related to infinitesimal shift interventions requiring only local overlap assumptions. ADEs, however, are rarely used in practice because their estimation usually requires estimating conditional density functions. By characterising the Riesz representers of weighted ADEs, we propose a new class of estimands that provides a unified view of weighted ADEs/ATEs when the treatment is continuous/binary. We derive the estimand in our class that minimises the nonparametric efficiency bound, thereby extending optimal weighting results from the binary treatment literature to the continuous setting. We develop efficient estimators for two weighted ADEs that avoid density estimation and are amenable to modern machine learning methods, which we evaluate in simulations and an applied analysis of Warfarin dosage effects.

math.ST

On estimands in target trial emulation

The target trial framework enables causal inference from longitudinal observational data by emulating randomized trials initiated at multiple time points. Precision is often improved by pooling information across trials, with standard models typically assuming - among other things - a time-constant treatment effect. However, this obscures interpretation when the true treatment effect varies, which we argue to be likely as a result of relying on noncollapsible estimands. To address these challenges, this paper introduces a model-free strategy for target trial analysis, centered around the choice of the estimand, rather than model specification. This ensures that treatment effects remain clearly interpretable for well-defined populations even under model misspecification. We propose estimands suitable for different study designs, and develop accompanying G-computation and inverse probability weighted estimators. Applications on simulations and real data on antimicrobial de-escalation in an intensive care unit setting demonstrate the greater clarity and reliability of the proposed methodology over traditional techniques.

stat.ME

Variable importance measures for heterogeneous treatment effects

Motivated by applications in precision medicine and treatment effect heterogeneity, recent research has focused on estimating conditional average treatment effects (CATEs) using machine learning (ML). CATE estimates may represent complicated functions that provide little insight into the key drivers of heterogeneity. Therefore, we introduce nonparametric treatment effect variable importance measures (TE-VIMs), based on the mean-squared error (MSE) in predicting the individual treatment effect. More precisely, TE-VIMs represent the increase in MSE when variables are removed from the CATE conditioning set. We derive efficient TE-VIM estimators which can be used with any CATE estimation strategy and are amenable to ML estimation. We propose several strategies to calculate these VIMs (e.g. leave-one out, or keep-one in), using popular meta-learners for the CATE. We study the finite sample performance through a simulation study and illustrate their application using clinical trial data.

stat.ME

Causal machine learning methods and use of cross-fitting in settings with high-dimensional confounding

Observational epidemiological studies commonly seek to estimate the causal effect of an exposure on an outcome. Adjustment for potential confounding bias in modern studies is challenging due to the presence of high-dimensional confounding, which occurs when there are many confounders relative to sample size or complex relationships between continuous confounders and exposure and outcome. Doubly robust methods such as Augmented Inverse Probability Weighting (AIPW) and Targeted Maximum Likelihood Estimation (TMLE) have the potential to address these challenges, using data-adaptive approaches and cross-fitting, but despite recent advances limited evaluation and guidance are available on their implementation in realistic settings where high-dimensional confounding is present. Motivated by an early-life cohort study, we conducted an extensive simulation study to compare the relative performance of AIPW and TMLE using data-adaptive approaches for estimating the average causal effect (ACE). We evaluated the benefits of using cross-fitting with a varying number of folds, as well as the impact of using a reduced versus full (larger, more diverse) library in the Super Learner ensemble learning approach used for implementation. We found that AIPW and TMLE performed similarly in most cases for estimating the ACE, but TMLE was more stable. Cross-fitting improved the performance of both methods, but was more important for variance estimation and coverage than for point estimates, with the number of folds a less important consideration. Using a full Super Learner library was important to reduce bias and variance in complex scenarios typical of modern health research studies.

stat.ME

Propensity weighting plus adjustment in proportional hazards model is not doubly robust

Recently, it has become common for applied works to combine commonly used survival analysis modeling methods, such as the multivariable Cox model and propensity score weighting, with the intention of forming a doubly robust estimator of an exposure effect hazard ratio that is unbiased in large samples when either the Cox model or the propensity score model is correctly specified. This combination does not, in general, produce a doubly robust estimator, even after regression standardization, when there is truly a causal effect. We demonstrate via simulation this lack of double robustness for the semiparametric Cox model, the Weibull proportional hazards model, and a simple proportional hazards flexible parametric model, with both the latter models fit via maximum likelihood. We provide a novel proof that the combination of propensity score weighting and a proportional hazards survival model, fit either via full or partial likelihood, is consistent under the null of no causal effect of the exposure on the outcome under particular censoring mechanisms if either the propensity score or the outcome model is correctly specified and contains all confounders. Given our results suggesting that double robustness only exists under the null, we outline two simple alternative estimators that are doubly robust for the survival difference at a given time point (in the above sense), provided the censoring mechanism can be correctly modeled, and one doubly robust method of estimation for the full survival curve. We provide R code to use these estimators for estimation and inference in the supporting information.

stat.ME

Causal machine learning for high-dimensional mediation analysis using interventional effects mapped to a target trial

Causal mediation analysis examines causal pathways linking exposures to disease. The estimation of interventional effects, which are mediation estimands that overcome certain identifiability problems of natural effects, has been advanced through causal machine learning methods, particularly for high-dimensional mediators. Recently, it has been proposed interventional effects can be defined in each study by mapping to a target trial assessing specific hypothetical mediator interventions. This provides an appealing framework to directly address real-world research questions about the extent to which such interventions might mitigate an increased disease risk in the exposed. However, existing estimators for interventional effects mapped to a target trial rely on singly-robust parametric approaches, limiting their applicability in high-dimensional settings. Building upon recent developments in causal machine learning for interventional effects, we address this gap by developing causal machine learning estimators for three interventional effect estimands, defined by target trials assessing hypothetical interventions inducing distinct shifts in joint mediator distributions. These estimands are motivated by a case study within the Longitudinal Study of Australian Children, used for illustration, which assessed how intervening on high inflammatory burden and other non-inflammatory adverse metabolomic markers might mitigate the adverse causal effect of overweight or obesity on high blood pressure in adolescence. We develop one-step and (partial) targeted minimum loss-based estimators based on efficient influence functions of those estimands, demonstrating they are root-n consistent, efficient, and multiply robust under certain conditions.

stat.ME

Causal machine learning for heterogeneous treatment effects in the presence of missing outcome data

When estimating heterogeneous treatment effects, missing outcome data can complicate treatment effect estimation, causing certain subgroups of the population to be poorly represented. In this work, we discuss this commonly overlooked problem and consider the impact that missing at random (MAR) outcome data has on causal machine learning estimators for the conditional average treatment effect (CATE). We propose two de-biased machine learning estimators for the CATE, the mDR-learner and mEP-learner, which address the issue of under-representation by integrating inverse probability of censoring weights into the DR-learner and EP-learner respectively. We show that under reasonable conditions, these estimators are oracle efficient, and illustrate their favorable performance through simulated data settings, comparing them to existing CATE estimators, including comparison to estimators which use common missing data techniques. We present an example of their application using the GBSG2 trial, exploring treatment effect heterogeneity when comparing hormonal therapies to non-hormonal therapies among breast cancer patients post surgery, and offer guidance on the decisions a practitioner must make when implementing these estimators.

stat.ML

Integration of aggregated data in causally interpretable meta-analysis by inverse weighting

Obtaining causally interpretable meta-analysis results is challenging when there are differences in the distribution of effect modifiers between eligible trials. To overcome this, recent work on transportability methods has considered standardizing results of individual studies over the case-mix of a target population, prior to pooling them as in a classical random-effect meta-analysis. One practical challenge, however, is that case-mix standardization often requires individual participant data (IPD) on outcome, treatments and case-mix characteristics to be fully accessible in every eligible study, along with IPD case-mix characteristics for a random sample from the target population. In this paper, we aim to develop novel strategies to integrate aggregated-level data from eligible trials with non-accessible IPD into a causal meta-analysis, by extending moment-based methods frequently used for population-adjusted indirect comparison in health technology assessment. Since valid inference for these moment-based methods by M-estimation theory requires additional aggregated data that are often unavailable in practice, computational methods to address this concern are also developed. We assess the finite-sample performance of the proposed approaches by simulated data, and then apply these on real-world clinical data to investigate the effectiveness of risankizumab versus ustekinumab among patients with moderate to severe psoriasis.

stat.ME

Debiasing Synthetic Data Generated by Deep Generative Models

While synthetic data hold great promise for privacy protection, their statistical analysis poses significant challenges that necessitate innovative solutions. The use of deep generative models (DGMs) for synthetic data generation is known to induce considerable bias and imprecision into synthetic data analyses, compromising their inferential utility as opposed to original data analyses. This bias and uncertainty can be substantial enough to impede statistical convergence rates, even in seemingly straightforward analyses like mean calculation. The standard errors of such estimators then exhibit slower shrinkage with sample size than the typical 1 over root-$n$ rate. This complicates fundamental calculations like p-values and confidence intervals, with no straightforward remedy currently available. In response to these challenges, we propose a new strategy that targets synthetic data created by DGMs for specific data analyses. Drawing insights from debiased and targeted machine learning, our approach accounts for biases, enhances convergence rates, and facilitates the calculation of estimators with easily approximated large sample variances. We exemplify our proposal through a simulation study on toy data and two case studies on real-world data, highlighting the importance of tailoring DGMs for targeted data analysis. This debiasing strategy contributes to advancing the reliability and applicability of synthetic data in statistical inference.

stat.ML

On Doubly Robust Inference for Double Machine Learning in Semiparametric Regression

Due to concerns about parametric model misspecification, there is interest in using machine learning to adjust for confounding when evaluating the causal effect of an exposure on an outcome. Unfortunately, exposure effect estimators that rely on machine learning predictions are generally subject to so-called plug-in bias, which can render naive $p$-values and confidence intervals invalid. Progress has been made via proposals like targeted minimum loss estimation and more recently double machine learning, which rely on learning the conditional mean of both the outcome and exposure. Valid inference can then be obtained so long as both predictions converge (sufficiently fast) to the truth. Focusing on partially linear regression models, we show that a specific implementation of the machine learning techniques can yield exposure effect estimators that have small bias even when one of the first-stage predictions does not converge to the truth. The resulting tests and confidence intervals are doubly robust. We also show that the proposed estimators may fail to be regular when only one nuisance parameter is consistently estimated; nevertheless, we observe in simulation studies that our proposal can lead to reduced bias and improved confidence interval coverage in moderate-to-large samples.

stat.ME

Chasing Shadows: How Implausible Assumptions Skew Our Understanding of Causal Estimands

The ICH E9 (R1) addendum on estimands, coupled with recent advancements in causal inference, has prompted a shift towards using model-free treatment effect estimands that are more closely aligned with the underlying scientific question. This represents a departure from traditional, model-dependent approaches where the statistical model often overshadows the inquiry itself. While this shift is a positive development, it has unintentionally led to the prioritization of an estimand's ability to perfectly answer the key scientific question over its practical learnability from data under plausible assumptions. We illustrate this by scrutinizing assumptions in the recent clinical trials literature on principal stratum estimands, demonstrating that some popular assumptions are not only implausible but often inevitably violated. We advocate for a more balanced approach to estimand formulation, one that carefully considers both the scientific relevance and the practical feasibility of estimation under realistic conditions.

stat.ME

Two Stage Least Squares with Time-Varying Instruments: An Application to an Evaluation of Treatment Intensification for Type-2 Diabetes

As longitudinal data becomes more available in many settings, policy makers are increasingly interested in the effect of time-varying treatments (e.g. sustained treatment strategies). In settings such as this, the preferred analysis techniques are the g-methods, however these require the untestable assumption of no unmeasured confounding. Instrumental variable analyses can minimise bias through unmeasured confounding. Of these methods, the Two Stage Least Squares technique is one of the most well used in Econometrics, but it has not been fully extended, and evaluated, in full time-varying settings. This paper proposes a robust two stage least squares method for the econometric evaluation of time-varying treatment. Using a simulation study we found that, unlike standard two stage least squares, it performs relatively well across a wide range of circumstances, including model misspecification. It compares well with recent time-varying instrument approaches via g-estimation. We illustrate the methods in an evaluation of treatment intensification for Type-2 Diabetes Mellitus, exploring the exogeneity in prescribing preferences to operationalise a time-varying instrument.

stat.ME

Instrumental Variable methods to target Hypothetical Estimands with longitudinal repeated measures data: Application to the STEP 1 trial

The STEP 1 randomized trial evaluated the effect of taking semaglutide vs placebo on body weight over a 68 week duration. As with any study evaluating an intervention delivered over a sustained period, non-adherence was observed. This was addressed in the original trial analysis within the Estimand Framework by viewing non-adherence as an intercurrent event. The primary analysis applied a treatment policy strategy which viewed it as an aspect of the treatment regimen, and thus made no adjustment for its presence. A supplementary analysis used a hypothetical strategy, targeting an estimand that would have been realised had all participants adhered, under the assumption that no post-baseline variables confounded adherence and change in body weight. In this paper we propose an alternative Instrumental Variable method to adjust for non-adherence which does not rely on the same `unconfoundedness' assumption and is less vulnerable to positivity violations (e.g., it can give valid results even under conditions where non-adherence is guaranteed). Unlike many previous Instrumental Variable approaches, it makes full use of the repeatedly measured outcome data, and allows for a time-varying effect of treatment adherence on a participant's weight. We show that it provides a natural vehicle for defining two distinct hypothetical estimands: the treatment effect if all participants would have adhered to semaglutide, and the treatment effect if all participants would have adhered to both semaglutide and placebo. When applied to the STEP 1 study, they both suggest a sustained, slowly decaying weight loss effect of semaglutide treatment.

stat.ME

The Real Deal Behind the Artificial Appeal: Inferential Utility of Tabular Synthetic Data

Recent advances in generative models facilitate the creation of synthetic data to be made available for research in privacy-sensitive contexts. However, the analysis of synthetic data raises a unique set of methodological challenges. In this work, we highlight the importance of inferential utility and provide empirical evidence against naive inference from synthetic data, whereby synthetic data are treated as if they were actually observed. Before publishing synthetic data, it is essential to develop statistical inference tools for such data. By means of a simulation study, we show that the rate of false-positive findings (type 1 error) will be unacceptably high, even when the estimates are unbiased. Despite the use of a previously proposed correction factor, this problem persists for deep generative models, in part due to slower convergence of estimators and resulting underestimation of the true standard error. We further demonstrate our findings through a case study.

cs.LG