arXiv ScienceSearch

arXiv subjects

Mathias Drton

Publications and source records attributed to Mathias Drton.

At least 19 recordsLinked to original sources

Goodness-of-Fit Tests for Linear Non-Gaussian Structural Equation Models

The field of causal discovery develops model selection methods to infer cause-effect relations among a set of random variables. For this purpose, different modelling assumptions have been proposed to render cause-effect relations identifiable. One prominent assumption is that the joint distribution of the observed variables follows a linear non-Gaussian structural equation model. In this paper, we develop novel goodness-of-fit tests that assess the validity of this assumption in the basic setting without latent confounders as well as in extension to linear models that incorporate latent confounders. Our approach involves testing algebraic relations among second and higher moments that hold as a consequence of the linearity of the structural equations. Specifically, we show that the linearity implies rank constraints on matrices and tensors derived from moments. For a practical implementation of our tests, we consider a multiplier bootstrap method that uses incomplete U-statistics to estimate subdeterminants, as well as asymptotic approximations to the null distribution of singular values. The methods are illustrated, in particular, for the Tübingen collection of benchmark data sets on cause-effect pairs.

stat.ME

Asymptotics for Model Selection in Probabilistic Principal Component Analysis

The probabilistic formulation of principal component analysis promises statistically grounded solutions to the problem of selecting the number of principal components. However, developing tractable model selection methods is complicated by the fact that the probabilistic principal component analysis (PPCA) model exhibits non-standard large sample asymptotics at model singularities, where the Fisher information matrix does not have full rank. In this work, tools from singular learning theory are used to provide a complete description of the marginal likelihood asymptotics for PPCA. The singular Bayesian information criterion (sBIC) along with our asymptotic results provides an effective procedure for selecting the number of principal components. In particular, the sBIC corrects for the overpenalization of the standard dimension-based BIC, which leads to too few principal components being selected, while also selecting the smallest true model with probability converging to one. Our framework extends beyond PPCA, where we also provide expressions for the sBIC that can be used to select between PPCA and factor analysis models. In both simulations and on real data the effectiveness of the proposed sBIC methodology is demonstrated.

math.ST

Semiparametric Inference for Half-Trek Estimators in Linear Structural Equation Models

Linear structural equation models on directed mixed graphs encode causal relationships among variables subject to latent confounding. The half-trek criterion (HTC) provides a graphical sufficient condition for the structural coefficients to be rationally identifiable from the observable covariance matrix, and yields a corresponding closed-form rational estimator. Despite this, the asymptotic distribution of the HTC estimator, and hence valid standard errors and confidence regions, have not been derived. We derive the semiparametric influence function of this estimator for all HTC-identified directed mixed graphs, including cyclic ones. The influence function combines the structural residual at the target node with the identification instruments, recursively corrected for uncertainty from earlier estimation stages. The HTC estimator is asymptotically normal with variance computable in closed form, yielding confidence regions, marginal intervals, and Wald tests for individual structural coefficients. Applied to the Fulton Fish Market dataset, our theory delivers a complete inferential summary for the causal effect of supply on demand.

stat.ME

Non-parametric recovery of causal diffusion mechanisms from steady-state observations

We consider sparse multivariate stochastic systems that evolve in continuous time according to a causal mechanism and present methodology to recover the system's time-infinitesimal transition mechanism from mere cross-sectional data. This observational paradigm is motivated by applications such as gene expression analysis, where destructive experimental techniques may only allow recording data once over a cell's lifetime. Precisely, we assume the system follows a time-homogeneous diffusion process that has reached an equilibrium distribution at observation time. Further, we assume the causal mechanism is fully described by the diffusion drift, is acyclic, and its causal structure graph is known. In this setting, we prove that the full causal mechanism, i.e., the drift function, can be non-parametrically identified under a weak non-explosion criterion. We derive a non-parametric kernel estimator for this challenging inverse problem and prove its consistency. Moreover, we propose a cross-validation scheme for hyperparameter tuning, illustrate the behavior of our estimator in simulations, and we discuss connections with irreversible generative diffusion models and low-frequency sampled data.

stat.ML

Singular Learning Theory for Factor Analysis

Watanabe's singular learning theory provides a framework for asymptotic analysis of Bayesian model selection for statistical models with singularities, where traditional statistical regularity assumptions fail. Learning coefficients, also known as real log canonical thresholds, play a central role in singular learning, as they govern the asymptotic behavior of Bayesian marginal likelihood integrals in settings where the Laplace approximations used for regular statistical models are not applicable. Learning coefficients are algebraic invariants that quantify the geometric complexity of a model and reveal how the singular structure impacts the model's generalization properties. In this paper, we apply algebraic methods to study the learning coefficients of factor analysis models, which are widely used latent variable models for continuously distributed data. Our main result provides exact formulas for learning coefficients of factor analysis models. Moreover, we study the singularity types of specific factor analysis models in detail.

math.ST

Mapping the causal structure of price formation in Texas's transitioning electricity market

Renewable deployment and rising demand from electrification and large digital loads are transforming electricity markets. However, how these developments reshape electricity price dynamics remains poorly understood, leaving system planners, capacity investors, and market participants reliant on assumptions from a thermal-dominated era that may no longer hold. We use causal discovery to study the evolution of wholesale electricity prices in Texas, which is undergoing rapid transformation. Our findings overturn the view of Texas as a gas-price-driven market, demonstrating that wind generation has become the dominant causal driver of day-ahead prices, with effects more than three times greater than those of natural gas. Yet wind's price-suppressing effect is weakening during peak periods, and wind growth redistributes congestion costs to distant load centres. Furthermore, rising load in South and West Texas alters system prices and regional differentials. Uncovering the evolving spatiotemporal nature of causal drivers, our analysis reveals that the pace, geographic siting, and relative scale of new generation and large loads will be decisive for future electricity price risks, infrastructure needs, and investments.

econ.GN

Identifying Direct Causal Effects in Latent Factor Models by Accounting for Unidentified Parents

We consider linear structural equation models with explicitly modelled latent variables. In such models, observed and latent variables solve linear equations including stochastic noise terms. The goal of our work is to identify the direct causal effects between the observed variables of interest by providing (rational) formulas in the observed covariances. Most prior identification approaches operate in the latent projection framework, where latent variables are projected away into dependent error terms. However, when the observed variables are densely confounded, even if only by a few latent variables, the projection-based approaches are unable to certify identifiability of most effects. For such problems, approaches that explicitly use the latent variables are more effective, but algorithms that were recently proposed for this purpose often remain inconclusive for denser causal graphs. We develop a new identification criterion that is able to better handle dense graphs by leveraging the key insight that recursive identification schemes can be generalized by explicitly accounting for causal parents with (yet) unidentified direct effects. Combinatorial search problems in our new criterion can be tackled with the help of network-flow computations, leading to a practical useful algorithmic tool that we also make available in software.

stat.ME

Cost-Aware Optimized Front-Door Experimental Design

Causal effect estimation often succeeds cost-constrained sequential data collection. This work considers multivariate linear front-door models with arbitrary unobserved confounding on treatment and response. We optimize the experimental design by balancing the statistical efficiency and measurement costs through partial data. The full-data efficient influence function for the causal effect is derived, together with the geometry of all observed-data influence functions. This characterization yields a closed-form optimal sampling policy and an estimator to minimize the asymptotic variance of regular asymptotically linear (RAL) estimators within a class of augmented full-data influence functions. The resulting design also covers back-door estimation. In simulations and applications to biological, medical, and industrial datasets, the optimized designs achieve substantial efficiency gains ($5.3\%$ to $31.9\%$) over naive full-sampling strategies.

stat.ME

Parameter identification in linear non-Gaussian causal models under general confounding

Linear non-Gaussian causal models postulate that each random variable is a linear function of parent variables and non-Gaussian exogenous error terms. We study identification of the linear coefficients when such models contain latent variables. Our focus is on the commonly studied acyclic setting, where each model corresponds to a directed acyclic graph (DAG). For this case, prior literature has demonstrated that connections to overcomplete independent component analysis yield effective criteria to decide parameter identifiability in latent variable models. However, this connection is based on the assumption that the observed variables linearly depend on the latent variables. Departing from this assumption, we treat models that allow for arbitrary non-linear latent confounding. Our main result is a graphical criterion that is necessary and sufficient for deciding the generic identifiability of direct causal effects. Moreover, we provide an algorithmic implementation of the criterion with a run time that is polynomial in the number of observed variables. Finally, we report on estimation heuristics based on the identification result and explore a generalization to models with feedback loops.

stat.ME

On the distance between mean and geometric median in high dimensions

The geometric median, a notion of center for multivariate distributions, has gained recent attention in robust statistics and machine learning. Although conceptually distinct from the mean (i.e., expectation), we demonstrate that both are very close in high dimensions when the dependence between the distribution components is suitably controlled. Concretely, we find an upper bound on the distance that vanishes with the dimension asymptotically, and derive a rate-matching first order expansion of the geometric median components. Simulations illustrate and confirm our results.

math.ST

Efficient Learning of Stationary Diffusions with Stein-type Discrepancies

Learning a stationary diffusion amounts to estimating the parameters of a stochastic differential equation whose stationary distribution matches a target distribution. We build on the recently introduced kernel deviation from stationarity (KDS), which enforces stationarity by evaluating expectations of the diffusion's generator in a reproducing kernel Hilbert space. Leveraging the connection between KDS and Stein discrepancies, we introduce the Stein-type KDS (SKDS) as an alternative formulation. We prove that a vanishing SKDS guarantees alignment of the learned diffusion's stationary distribution with the target. Furthermore, under broad parametrizations, SKDS is convex with an empirical version that is $ε$-quasiconvex with high probability. Empirically, learning with SKDS attains comparable accuracy to KDS while substantially reducing computational cost and yields improvements over the majority of competitive baselines.

stat.ML

Matching Criterion for Identifiability in Sparse Factor Analysis

Factor analysis models explain dependence among observed variables by a smaller number of unobserved factors. A main challenge in confirmatory factor analysis is determining whether the factor loading matrix is identifiable from the observed covariance matrix. The factor loading matrix captures the linear effects of the factors and, if unrestricted, can only be identified up to an orthogonal transformation of the factors. However, in many applications the factor loadings exhibit an interesting sparsity pattern that may lead to identifiability up to column signs. We study this phenomenon by connecting sparse confirmatory factor analysis models to bipartite graphs and providing sufficient graphical conditions for identifiability of the factor loading matrix up to column signs. In contrast to previous work, our main contribution, the matching criterion, exploits sparsity by operating locally on the graph structure, thereby improving existing conditions. Our criterion is efficiently decidable in time that is polynomial in the size of the graph, when restricting the search steps to sets of bounded size.

math.ST

On weak convergence of Gaussian conditional distributions

Weak convergence of joint distributions generally does not imply convergence of conditional distributions. In particular, conditional distributions need not converge when joint Gaussian distributions converge to a singular Gaussian limit. Algebraically, this is due to the fact that at singular covariance matrices, Schur complements are not continuous functions of the matrix entries. Our results lay out special conditions under which convergence of Gaussian conditional distributions nevertheless occurs, and we exemplify how this allows one to reason about conditional independence in a new class of graphical models.

math.ST

Trek-Based Parameter Identification for Linear Causal Models With Arbitrarily Structured Latent Variables

We develop a criterion to certify whether causal effects are identifiable in linear structural equation models with latent variables. Linear structural equation models correspond to directed graphs whose nodes represent the random variables of interest and whose edges are weighted with linear coefficients that correspond to direct causal effects. In contrast to previous identification methods, we do not restrict ourselves to settings where the latent variables constitute independent latent factors (i.e., to source nodes in the graphical representation of the model). Our novel latent-subgraph criterion is a purely graphical condition that is sufficient for identifiability of causal effects by rational formulas in the covariance matrix. To check the latent-subgraph criterion, we provide a sound and complete algorithm that operates by solving an integer linear program. While it targets effects involving observed variables, our new criterion is also useful for identifying effects between latent variables, as it allows one to transform the given model into a simpler measurement model for which other existing tools become applicable.

math.ST

Causal Discovery for Linear Non-Gaussian Models with Disjoint Cycles

The paradigm of linear structural equation modeling readily allows one to incorporate causal feedback loops in the model specification. These appear as directed cycles in the common graphical representation of the models. However, the presence of cycles entails difficulties such as the fact that models need no longer be characterized by conditional independence relations. As a result, learning cyclic causal structures remains a challenging problem. In this paper, we offer new insights on this problem in the context of linear non-Gaussian models. First, we precisely characterize when two directed graphs determine the same linear non-Gaussian model. Next, we take up a setting of cycle-disjoint graphs, for which we are able to show that simple quadratic and cubic polynomial relations among low-order moments of a non-Gaussian distribution allow one to locate source cycles. Complementing this with a strategy of decorrelating cycles and multivariate regression allows one to infer a block-topological order among the directed cycles, which leads to a {consistent and computationally efficient algorithm} for learning causal structures with disjoint cycles.

math.ST

Causal Effect Identification in lvLiNGAM from Higher-Order Cumulants

This paper investigates causal effect identification in latent variable Linear Non-Gaussian Acyclic Models (lvLiNGAM) using higher-order cumulants, addressing two prominent setups that are challenging in the presence of latent confounding: (1) a single proxy variable that may causally influence the treatment and (2) underspecified instrumental variable cases where fewer instruments exist than treatments. We prove that causal effects are identifiable with a single proxy or instrument and provide corresponding estimation methods. Experimental results demonstrate the accuracy and robustness of our approaches compared to existing methods, advancing the theoretical and practical understanding of causal inference in linear systems with latent confounders.

stat.ML

Nonlinear Causal Discovery for Grouped Data

Inferring cause-effect relationships from observational data has gained significant attention in recent years, but most methods are limited to scalar random variables. In many important domains, including neuroscience, psychology, social science, and industrial manufacturing, the causal units of interest are groups of variables rather than individual scalar measurements. Motivated by these applications, we extend nonlinear additive noise models to handle random vectors, establishing a two-step approach for causal graph learning: First, infer the causal order among random vectors. Second, perform model selection to identify the best graph consistent with this order. We introduce effective and novel solutions for both steps in the vector case, demonstrating strong performance in simulations. Finally, we apply our method to real-world assembly line data with partial knowledge of causal ordering among variable groups.

stat.ML

On universal inference in Gaussian mixture models

A recent line of work provides new statistical tools based on game-theory and achieves safe anytime-valid inference without assuming regularity conditions. In particular, the framework of universal inference proposed by Wasserman, Ramdas and Balakrishnan [78] offers new solutions to testing problems by modifying the likelihood ratio test in a data-splitting scheme. In this paper, we study the performance of the resulting split likelihood ratio test under Gaussian mixture models, which are canonical examples for models in which classical regularity conditions fail to hold. We establish that under the null hypothesis, the split likelihood ratio statistic is asymptotically normal with increasing mean and variance. Contradicting the usual belief that the flexibility of universal inference comes at the price of a significant loss of power, we prove that universal inference surprisingly achieves the same detection rate $(n^{-1}\log\log n)^{1/2}$ as the classical likelihood ratio test.

math.ST