arXiv · 2405.19225
Synthetic Potential Outcomes and Causal Mixture Identifiability
Abstract
Heterogeneous data from multiple populations, sub-groups, or sources is often represented as a ``mixture model'' with a single latent class influencing all of the observed covariates. Heterogeneity can be resolved at multiple levels by grouping populations according to different notions of similarity. This paper proposes grouping with respect to the causal response of an intervention or perturbation on the system. This definition is distinct from previous notions, such as similar covariate values (e.g. clustering) or similar correlations between covariates (e.g. Gaussian mixture models). To solve the problem, we ``synthetically sample'' from a counterfactual distribution using higher-order multi-linear moments of the observable data. To understand how these ``causal mixtures'' fit in with more classical notions, we develop a hierarchy of mixture identifiability.
Explore related subjects
Keep this discovery
Bijan Mazaheri, Chandler Squires, Caroline Uhler. 2024-05-29. Synthetic Potential Outcomes and Causal Mixture Identifiability. https://arxiv.org/abs/2405.19225
Cite the original work for its findings. Save a collection to share your selection of sources.