arXiv Science⌕ Search

arXiv · 2609.28856

Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions

Abstract

Causal discovery from observational and interventional data becomes challenging in the presence of latent confounding and selection bias, where causal structure is no longer adequately represented by directed acyclic graphs over observed variables. Existing model-free methods often rely on an exponential number of conditional independence tests and provide limited uncertainty quantification in high-dimensional settings. We develop a model-free and constraint-query optimal statistical inference framework for causal discovery under latent variables and selection using single-target interventions. We introduce the system-induced subgraph (SIS) to capture the causal relations among system variables while accounting for context variables. We establish its identifiability through maximal ancestral graphs (MAGs), and show that interventions on each observed system variable are sufficient for unique identification and necessary in the worst case. Building on these results, we develop a two-stage graph inference procedure with asymptotic family-wise error control under sufficient first-stage power. For $d_X$ observed system variables, the procedure requires at most $\frac{5}{2}d_X^2$ statistical tests, parallelizable within each stage, and achieves optimal constraint-query complexity up to a constant factor. The framework accommodates soft interventions and avoids parametric structural equation assumptions. We illustrate the methods through analysis of Perturb-seq data from interferon-$β$-stimulated A549 lung cancer cell lines.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xiaotian Hou, Kwangmoon Park, Hongzhe Li. 2026-09-23. Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions. https://arxiv.org/abs/2609.28856

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PK/PD-integrated Bayesian platform design for phase II dose regimen optimization

Early-phase dose-finding methods increasingly assess toxicity and efficacy jointly, but comparisons based only on administered dose may inadequately characterize regimens differing in schedule. We developed a Bayesian phase II adaptive platform design for regimen optimization that integrates pharmacokinetic/pharmacodynamic (PK/PD) modelling into toxicity, efficacy, regimen selection and adaptation decisions. The proposed PK/PD-informed Regimen Optimization Platform (PROP) design uses a population PK/PD model to generate patient- and population-level predictions of exposure and biological activity. Acute and cumulative toxicities are analysed using a discrete-time time-to-event model informed by PK exposure. Efficacy is evaluated through Bayesian model averaging of exposure-driven and biomarker-driven time-to-event models. The design supports regimen graduation, discontinuation for futility or safety, and addition of unexplored regimens. Performance was evaluated through simulations motivated by an influenza intensive-care setting. Across six scenarios, PROP generally improved graduation and futility decisions, reduced inappropriate graduation, and supported the addition of promising regimens compared with dose-based alternatives. It also more accurately estimated regimen-specific toxicity and arm-specific efficacy, while the model-averaging framework favored the efficacy model consistent with the data-generating mechanism. Dose-based approaches performed better for safety stopping in some scenarios, despite less accurate characterization of the regimen--toxicity relationship. PK/PD-informed platform designs can improve adaptive regimen selection and knowledge generation when dose alone cannot adequately characterize treatment regimens.

stat.ME↗

Towards more plausible point-identifying assumptions in two-sample Mendelian randomization

Two-sample Mendelian randomization (MR) is a widely applied methodology in epidemiology. In two-sample MR, summary data (typically, regression coefficients and standard errors) quantifying the association between multiple genetic variants and the exposure and the outcome are used in an instrumental variable framework aimed at estimating the causal effect of the exposure on the outcome. Most two-sample MR methods were developed under data-generating models where the association of for each candidate genetic instrument with the exposure, as well as the causal effect of the exposure on the outcome, are constant in the additive scale. These assumptions are useful because they imply that, had all genetic variants been valid IVs, they would all estimate the same causal parameter - namely, the constant causal effect. We refer to this condition as summary-level homogeneity. However, these are rather strong homogeneity conditions which may raise concerns about the plausibility of these methods in practice. In this paper, we show that summary-level homogeneity is implied by the following conditions: the causal effect is additive linear, but not necessarily constant across, all strata of the population; and uncorrelatedness between heterogeneity in the causal effect and in the association between each genetic variant and the exposure. Under these conditions, typical two-sample MR methods can be interpreted as estimators of the average causal effect. These results clarify that point-identifying assumptions required for two-sample MR methods are weaker than previously anticipated, which contributes to their plausibility and interpretation in at least some practical applications.

stat.ME↗

Amortized Bayesian Disease Mapping and Boundary Detection on Heterogeneous Spatial Graphs

Spatial disease maps help public-health researchers identify geographic inequalities, but standard Bayesian smoothing can obscure localized disparities when neighboring communities have sharply different socioeconomic or behavioral profiles. Analysts therefore need to determine where smoothing should be interrupted and repeat that analysis as maps, adjacency structures, and outcomes change. We develop a covariate-informed Bayesian boundary model and an amortized posterior approximation trained across heterogeneous areal graphs. The model distinguishes local interruptions in smoothing from broader residual spatial dependence; the trained approximation handles maps with different numbers of regions. Simulations examine posterior calibration, boundary-probability recovery, replicated-data behavior, and MCSE-controlled agreement with prior-matched MCMC. In contrast to traditional approaches that analyze these data separately, we demonstrate the effectiveness of using a single trained deep learning network to analyze respiratory hospitalizations in Greater Glasgow; lung cancer incidence in California; and tracheal, bronchial, and lung cancer mortality in South Korea, comprising 58 to 241 regions. Selected boundary density is greatest in Glasgow and lowest in South Korea despite substantial residual spatial dependence in both, showing that local interruption and broader spatial persistence need not vary together. Across all three applications, edge-level boundary probabilities agree substantially with dataset-specific analyses, although posterior spread and thresholded boundary sets differ. These results support reusable Bayesian boundary analysis across the evaluated disease-map class and identify the validation needed before deployment to new applications.

stat.ME↗