arXiv Science⌕ Search

arXiv · 2610.09100

Score-Based Methods for Selecting Direct Causes in Reweighted Distributions

Abstract

Scoring methods based on the Bayesian information criterion (BIC) are commonly used for model selection tasks such as choosing features in a regression model or learning directed acyclic graphs from data. In certain model classes, model selection based on the BIC is consistent. However, since the BIC is based on the observed likelihood function, it does not apply directly to reweighted distributions arising in causal inference, missing data, and domain shift applications. Here, we propose a generalized version of the BIC that allows for model selection in reweighted distributions. We prove its corresponding consistency property and demonstrate how it can be used for selecting direct causes of an outcome variable in a marginal structural model. Our proposed method accounts for scenarios with unmeasured confounders and where the weights must be estimated from data. Through simulation studies, we demonstrate the asymptotic properties of the reweighted BIC and compare it with alternative methods for model selection.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jacob M. Chen, Ilya Shpitser. 2026-10-06. Score-Based Methods for Selecting Direct Causes in Reweighted Distributions. https://arxiv.org/abs/2610.09100

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Robust Bayesian methods using amortized simulation-based inference

Bayesian simulation-based inference (SBI) methods are used in statistical models where simulation is feasible but the likelihood is intractable. Standard SBI methods can perform poorly in cases of model misspecification, and there has been much recent work on modified SBI approaches which are robust to misspecified likelihoods. However, less attention has been given to the issue of inappropriate prior specification, which is the focus of this work. In conventional Bayesian modelling, there will often be a wide range of prior distributions consistent with limited prior knowledge expressed by an expert. Choosing a single prior can lead to an inappropriate choice, possibly conflicting with the likelihood information. Robust Bayesian methods, where a class of priors is considered instead of a single prior, can address this issue. For each density in the prior class, a posterior can be computed, and the range of the resulting inferences is informative about posterior sensitivity to the prior imprecision. We consider density ratio classes for the prior and implement robust Bayesian SBI using amortized neural methods developed recently in the literature. We also discuss methods for checking for conflict between a density ratio class of priors and the likelihood, and sequential updating methods for examining conflict between different groups of summary statistics. The methods are illustrated for several simulated and real examples.

stat.ME↗

On the permutation equivariance principle for causal estimands

In many causal inference problems, multiple action variables share a common causal role yet lack a natural ordering. \revblue{We consider $K\geq2$ action variables, each evaluated under treatment or control conditions,} and formalize permutation equivariance, the principle that permuting the variables permutes the corresponding estimands in a trackable manner, hence preserving their scientific meaning. We characterize this principle algebraically and present a complete class of weighted permutation equivariant estimands capturing main effects and interactions of all orders. We discuss the interpretation and choice of weights and characterize residual-free estimands, whose inclusion--exclusion sum recovers the endpoint contrast between the all-treated and all-control configurations. \revblue{Applying our general framework to network interference yields a new hierarchy of direct, indirect, and overall effects, whose first-order aggregates recover the average effects of \citet{hu2022average}. We also identify the overall effects as mixed derivatives of expected welfare under independent Bernoulli assignment.} We illustrate the framework through the contexts of factorial studies, causal mediation, and network interference.

stat.ME↗

Testing the Deviation of a High-dimensional Mean

This paper investigates the high-dimensional one-sample equivalence testing problem for a reference mean vector $\bbmu_0$: $H_0:\|\bbmu-\bbmu_0\|>d_0$ versus $H_1: \|\bbmu-\bbmu_0\|\leq d_0$ for a pre-specified length threshold $d_0 > 0$. In high-dimensional regimes, classical equivalence tests suffer from severe asymptotic breakdowns. To resolve this difficulty, we propose a novel test statistic based on a two-armed bandit (TAB) U-process, which leverages positive/negative feedback mechanisms from control theory. We rigorously establish the weak convergence of TAB U-process to a novel two-dimensional stochastic differential equation (SDE), termed the U-bandit SDE. By employing in-depth analysis of this SDE, we characterize the asymptotic size and power of the resulting test statistic. The theoretical framework is extended to the two-sample setting and validated through finite-sample simulations and an empirical analysis of microbiome compositions.

stat.ME↗