arXiv ScienceSearch

arXiv subjects

Tuo Lin

Publications and source records attributed to Tuo Lin.

10 recordsLinked to original sources

Win-Ratio Regression for Prioritized Composite Outcomes in Observational Studies: Doubly Robust and Efficient Estimation with Future-Score Correction

Prioritized pairwise outcomes are useful when clinical events follow a natural hierarchy, but censoring before pair resolution complicates estimation. We develop a win-ratio regression framework for this setting by defining a complete-data target over follow-up and deriving an estimating equation for the observed data. The central idea is future-score correction (FC): when censoring prevents later pairwise comparisons from being observed, the method replaces the remaining score with its conditional expectation given the observed history. This correction recovers pairwise information beyond that provided by inverse censoring weights alone. Additionally, we incorporate treatment weighting and baseline outcome augmentation to address baseline confounding. Together, these components yield double robustness for treatment assignment and censoring. Inference is obtained from U-statistic theory. Under standard regularity conditions, the AIPW-FC estimator is asymptotically normal and efficient when all nuisance functions are correctly specified. Simulations with 30%, 50%, and 65% censoring show that efficiency gains from future-score correction increase with the censoring rate, with relative efficiency reaching 1.50 under 65% censoring and near-nominal coverage for AIPW-FC. An application to OneFlorida electronic health record data illustrates the method for a composite outcome that prioritizes death over hospitalization.

stat.ME

Multivariate incremental effects for continuous treatments: Studying the health effects of environmental mixtures

Evaluating the causal health effects of multivariate, continuous exposures, such as air pollution mixtures, is a critical public health challenge. A primary obstacle is the frequent violation of the positivity assumption, which renders the effects of standard deterministic interventions unidentified or heavily reliant on unreliable model extrapolation. In this paper, we develop a novel causal inference framework to address this challenge. We extend exponential tilting to multivariate exposures and address the critical question of how to compare different intervention directions fairly. This establishes a systematic framework for defining and evaluating various policy-relevant causal estimands, allowing researchers to address diverse scientific questions. We develop numerous methodological advancements, including efficient one-step estimation strategies, a Riemannian BFGS algorithm to solve a constrained manifold optimization problem, semiparametric efficiency bounds for causal estimands, minimax rates for estimators, and establishing asymptotic normality. We demonstrate our framework's utility by applying it to a nationwide environmental health dataset to identify the optimal strategy for reducing adverse health outcomes associated with a PM$_{2.5}$ chemical mixture.

stat.ME

Semiparametric Estimation of Delayed-Outcome Treatment Effects Using Short-Term Surrogates under Administrative Censoring

The multi-site registry studies, such as Stepped-wedge cluster-randomized trials (SW-CRT), staggered-enrollment RCTs, etc., share a structural feature: the primary long-term outcome is administratively censored for a non-negligible fraction of units, with censoring driven by calendar design rather than by the outcome itself. Standard inverse-probability-of-censoring weighting becomes unstable when observation probabilities $g_{\Delta}$ concentrate near zero for late-crossing units, while parametric mixed-model analyses discard the information in any short-term intermediate measurement and rely on correct specification of the secular time trend. We study semiparametric estimation of the average treatment effect when a short-term surrogate, which is observed for all units and conditionally independent of the censoring mechanism given baseline covariates, is available. Identification takes a nested-integral form in which the outcome regression is marginalized over the conditional surrogate distribution, so the observation mechanism does not enter the target functional as an inverse weight. We show that a density-plug-in one-step debiased machine-learning construction for this functional leaves a second-order cross-product remainder $R_{SY}$ that has no doubly-robust complement in the efficient influence function and is not eliminated by cross-fitting . We propose a surrogate-assisted AIPW estimator (SA-AIPW) that integrates over the empirical surrogate distribution through treatment weighting rather than estimating the conditional surrogate density, and so structurally avoids $R_{SY}$. For clustered data, the estimator is shown to be $\sqrt{J}$-consistent and asymptotically linear under a product-rate double-robustness condition.

stat.ME

Why Is the Double-Robust Estimator for Causal Inference Not Doubly Robust for Variance Estimation?

Doubly robust estimators (DRE) are widely used in causal inference because they yield consistent estimators of average causal effect when at least one of the nuisance models, the propensity for treatment (exposure) or the outcome regression, is correct. However, double robustness does not extend to variance estimation; the influence-function (IF)-based variance estimator is consistent only when both nuisance parameters are correct. This raises concerns about applying DRE in practice, where model misspecification is inevitable. The recent paper by Shook-Sa et al. (2025, Biometrics, 81(2), ujaf054) demonstrated through Monte Carlo simulations that the IF-based variance estimator is biased. However, the paper's findings are empirical. The key question remains: why does the variance estimator fail in double robustness, and under what conditions do alternatives succeed, such as the ones demonstrated in Shook-Sa et al. 2025. In this paper, we develop a formal theory to clarify the efficiency properties of DRE that underlie these empirical findings. We also introduce alternative strategies, including a mixture-based framework underlying the sample-splitting and crossfitting approaches, to achieve valid inference with misspecified nuisance parameters. Our considerations are illustrated with simulation and real study data.

stat.ME

Semiparametric Regression Models for Explanatory Variables with Missing Data due to Detection Limit

Detection limit (DL) has become an increasingly ubiquitous issue in statistical analyses of biomedical studies, such as cytokine, metabolite and protein analysis. In regression analysis, if an explanatory variable is left-censored due to concentrations below the DL, one may limit analyses to observed data. In many studies, additional, or surrogate, variables are available to model, and incorporating such auxiliary modeling information into the regression model can improve statistical power. Although methods have been developed along this line, almost all are limited to parametric models for both the regression and left-censored explanatory variable. While some recent work has considered semiparametric regression for the censored DL-effected explanatory variable, the regression of primary interest is still left parametric, which not only makes it prone to biased estimates, but also suffers from high computational cost and inefficiency due to maximizing an extremely complex likelihood function and bootstrap inference. In this paper, we propose a new approach by considering semiparametric generalized linear models (SPGLM) for the primary regression and parametric or semiparametric models for DL-effected explanatory variable. The semiparametric and semiparametric combination provides the most robust inference, while the semiparametric and parametric case enables more efficient inference. The proposed approach is also much easier to implement and allows for leveraging sample splitting and cross fitting (SSCF) to improve computational efficiency in variance estimation. In particular, our approach improves computational efficiency over bootstrap by 450 times. We use simulated and real study data to illustrate the approach.

stat.ME

Peak Inference for Gaussian Random Fields on a Lattice

In this work we develop a Monte Carlo method to compute the height distribution of local maxima of a stationary Gaussian or Gaussian-related random field that is observed on a regular lattice. We show that our method can be used to provide valid peak based inference in datasets with low levels of smoothness, where existing formulae derived for continuous domains are not accurate. We also extend the methods in Worsley (2005) and Taylor et al. (2007) to compute the peak height distribution and compare them with our approach. Lastly, we apply our method to a task fMRI dataset to show how it can be used in practice.

stat.ME

A doubly robust estimator for the Mann Whitney Wilcoxon Rank Sum Test when applied for causal inference in observational studies

The Mann-Whitney-Wilcoxon rank sum test (MWWRST) is a widely used method for comparing two treatment groups in randomized control trials, particularly when dealing with highly skewed data. However, when applied to observational study data, the MWWRST often yields invalid results for causal inference. To address this limitation, Wu et al. (2014) introduced an approach that incorporates inverse probability weighting (IPW) into this rank-based statistics to mitigate confounding effects. Subsequently, Mao (2018), Zhang et al. (2019), and Ai et al. (2020) extended this IPW estimator to develop doubly robust estimators. Nevertheless, each of these approaches has notable limitations. Mao's method imposes stringent assumptions that may not align with real-world study data. Zhang et al.'s (2019) estimators rely on bootstrap inference, which suffers from computational inefficiency and lacks known asymptotic properties. Meanwhile, Ai et al. (2020) primarily focus on testing the null hypothesis of equal distributions between two groups, which is a more stringent assumption that may not be well-suited to the primary practical application of MWWRST. In this paper, we aim to address these limitations by leveraging functional response models (FRM) to develop doubly robust estimators. We demonstrate the performance of our proposed approach using both simulated and real study data.

stat.ME

On Semiparametric Efficiency of an Emerging Class of Regression Models for Between-subject Attributes

The semiparametric regression models have attracted increasing attention owing to their robustness compared to their parametric counterparts. This paper discusses the efficiency bound for functional response models (FRM), an emerging class of semiparametric regression that serves as a timely solution for research questions involving pairwise observations. This new paradigm is especially appealing to reduce astronomical data dimensions for those arising from wearable devices and high-throughput technology, such as microbiome Beta-diversity, viral genetic linkage, single-cell RNA sequencing, etc. Despite the growing applications, the efficiency of their estimators has not been investigated carefully due to the extreme difficulty to address the inherent correlations among pairs. Leveraging the Hilbert-space-based semiparametric efficiency theory for classical within-subject attributes, this manuscript extends such asymptotic efficiency into the broader regression involving between-subject attributes and pinpoints the most efficient estimator, which leads to a sensitive signal-detection in practice. With pairwise outcomes burgeoning immensely as effective dimension-reduction summaries, the established theory will not only fill the critical gap in identifying the most efficient semiparametric estimator but also propel wide-ranging implementations of this new paradigm for between-subject attributes.

stat.ME

Estimating Viral Genetic Linkage Rates in the Presence of Missing Data

Although the interest in the the use of social and information networks has grown, most inferences on networks assume the data collected represents the complete. However, when ignoring missing data, even when missing completely at random, this results in bias for estimators regarding inference network related parameters. In this paper, we focus on constructing estimators for the probability that a randomly selected node has node has at least one edge under the assumption that nodes are missing completely at random along with their corresponding edges. In addition, issues also arise in obtaining asymptotic properties for such estimators, because linkage indicators across nodes are correlated preventing the direct application of the Central Limit Theorem and Law of Large Numbers. Using a subsampling approach, we present an improved estimator for our parameter of interest that accommodates for missing data. Utilizing the theory U-statistics, we derive consistency and asymptotic normality of the proposed estimator. This approach decreases the bias in estimating our parameter of interest. We illustrate our approach using the HIV viral strains from a large cluster-randomized trial of a combination HIV prevention intervention -- the Botswana Combination Prevention Project (BCPP).

stat.ME

A Riemann Manifold Model Framework for Longitudinal Changes in Physical Activity Patterns

Physical activity (PA) is significantly associated with many health outcomes. The wide usage of wearable accelerometer-based activity trackers in recent years has provided a unique opportunity for in-depth research on PA and its relations with health outcomes and interventions. Past analysis of activity tracker data relies heavily on aggregating minute-level PA records into day-level summary statistics, in which important information of PA temporal/diurnal patterns is lost. In this paper we propose a novel functional data analysis approach based on Riemann manifolds for modeling PA and its longitudinal changes. We model smoothed minute-level PA of a day as one-dimensional Riemann manifolds and longitudinal changes in PA in different visits as deformations between manifolds. The variability in changes of PA among a cohort of subjects is characterized via variability in the deformation. Functional principal component analysis is further adopted to model the deformations and PC scores are used as a proxy in modeling the relation between changes in PA and health outcomes and/or interventions. We conduct comprehensive analyses on data from two clinical trials: Reach for Health (RfH) and Metabolism, Exercise and Nutrition at UCSD (MENU), focusing on the effect of interventions on longitudinal changes in PA patterns and how different modes of changes in PA influence weight loss, respectively. The proposed approach reveals unique modes of changes including overall enhanced PA, boosted morning PA, and shifts of active hours specific to each study cohort. The results bring new insights into the study of longitudinal changes in PA and health and have the potential to facilitate designing of effective health interventions and guidelines.

stat.AP