arXiv ScienceSearch

arXiv subjects

Liam Solus

Publications and source records attributed to Liam Solus.

At least 19 recordsLinked to original sources

Algebraic implicitization techniques in the graphical models program

Many aspects of daily life now rely on technology driven by advanced techniques for analyzing multivariate data. Graphical models provide the general framework within which much of these analyses take place. However, different data types require different graphical models, with each requiring its own basic mathematical theory to guide sound statistical inference. Developing this theory for a family of graphical models requires solutions to fundamental questions, where emerging mathematical approaches are utilizing what can collectively be called algebraic implicitization techniques. These notes provide an introduction to the basics of graphical models and the algebraic implicitization techniques now appearing in instances of the graphical models program.

math.ST

Posets of trek polynomials for directed trees

When a variety $V_\varphi$ equals the image of a polynomial map $\varphi$ whose coordinate functions are combinatorial generating polynomials (i.e.~polynomials enumerating combinatorial objects), the geometry of $V_\varphi$ reflects identities satisfied by the generating polynomials. The resulting interplay between combinatorics and algebraic geometry can be used to answer questions about $V_\varphi$. A recent technique proposes to do so using a partially ordered set (poset) $P_\varphi$ defined via the coefficient vectors of the polynomials defining $\varphi$. This paper characterizes the poset $P_\varphi$ when the generating polynomials defining $\varphi$ enumerate subgraphs of a directed tree known as treks. The characterization is used to compute the linear span of $V_\varphi$, prove it is toric and deduce a basis for its vanishing ideal. It is also shown that this poset of trek polynomials for a directed tree is a so-called $\pi$-system if and only if the tree satisfies a property characterized via Stanley's P-partitions. As an additional consequence, it is shown that the varieties for two distinct directed trees intersect in a strictly lower-dimensional variety. This solves an instance of the structural identifiability problem in the graphical models program from statistics.

math.CO

On unirational varieties with poset parameterizations

We use partially ordered sets (posets) to provide a canonical parameterization for the Zariski closure of the image of a semialgebraic set under a rational map whose coordinate functions are polynomials with nonnegative integral coefficients. The resulting poset parametrization of such a unirational variety allows us to translate several well-studied problems into combinatorics; e.g. reducing the problems to describing the poset associated to the variety. These problems include, the implicitization problem from algebraic geometry, the toric reparameterization problem, the computation of the linear span of the variety, and the problem of distinguishing two semialgebraic subsets of the same ambient space. The technique applies to instances of these problems in several fields, including algebraic geometry, algebraic combinatorics, statistics and applied algebra. We demonstrate the technique on examples from each field, including degenerate subvarieties of secant varieties, matroid flat varieties -- which generalize toric varieties of edge polytopes, as well as varieties arising in multivariate data analysis and evolutionary biology.

math.CO

Intervening to Learn and Compose Causally Disentangled Representations

In designing generative models, it is commonly believed that in order to learn useful latent structure, we face a fundamental tension between expressivity and structure. In this paper we challenge this view by proposing a new approach to training arbitrarily expressive generative models that simultaneously learn causally disentangled concepts. This is accomplished by adding a simple context module to an arbitrarily complex black-box model, which learns to process concept information by implicitly inverting linear representations from the model's encoder. Inspired by the notion of intervention in a causal model, our module selectively modifies its architecture during training, allowing it to learn a compact joint model over different contexts. We show how adding this module leads to causally disentangled representations that can be composed for out-of-distribution generation on both real and simulated data. The resulting models can be trained end-to-end or fine-tuned from pre-trained models. To further validate our proposed approach, we prove a new identifiability result that extends existing work on identifying structured representations.

stat.ML

Ehrhart theory of cosmological polytopes

The cosmological polytope of a graph $G$ was recently introduced to give a geometric approach to the computation of wavefunctions for cosmological models with associated Feynman diagram $G$. Basic results in the theory of positive geometries dictate that this wavefunction may be computed as a sum of rational functions associated to the facets in a triangulation of the cosmological polytope. The normalized volume of the polytope then provides a complexity estimate for these computations. In this paper, we examine the (Ehrhart) $h^\ast$-polynomial of cosmological polytopes. We derive recursive formulas for computing the $h^\ast$-polynomial of disjoint unions and $1$-sums of graphs. The degree of the $h^\ast$-polynomial for any $G$ is computed and a characterization of palindromicity is given. Using these observations, a tight lower bound on the $h^\ast$-polynomial for any $G$ is identified and explicit formulas for the $h^\ast$-polynomials of multitrees and multicycles are derived. The results generalize the existing results on normalized volumes of cosmological polytopes. A tight upper bound and a combinatorial formula for the $h^\ast$-polynomial of any cosmological polytope are conjectured.

math.CO

Real birational implicitization for statistical models

We derive an implicit description of the image of a semialgebraic set under a birational map, provided that the denominators of the map are positive on the set. For statistical models which are globally rationally identifiable, this yields model-defining constraints which facilitate model membership testing, representation learning, and model equivalence tests. Many examples illustrate the applicability of our results. The implicit equations recover well-known Markov properties of classical graphical models, as well as other well-studied equations such as the Verma constraint. They also provide Markov properties for generalizations of these frameworks, such as colored or interventional graphical models, staged trees, and the recently introduced Lyapunov models. Under a further mild assumption, we show that our implicit equations generate the vanishing ideal of the model up to a saturation, generalizing previous results of Geiger, Meek and Sturmfels, Duarte and G\"orgen, Sullivant, and others.

math.ST

Colored Multiset Eulerian Polynomials

Colored multiset Eulerian polynomials are a common generalization of MacMahon's multiset Eulerian polynomials and the colored Eulerian polynomials, both of which are known to satisfy well-studied distributional properties including real-rootedness, log-concavity and unimodality. The symmetric colored multiset Eulerian polynomials are characterized and used to prove sufficient conditions for a colored multiset Eulerian polynomial to be self-interlacing. The latter property implies the aforementioned distributional properties as well as others, including the alternatingly increasing property and bi-$\gamma$-positivity. To derive these results, multivariate generalizations of an identity due to MacMahon are deduced. The results are applied to a pair of questions, both previously studied in several special cases, that are seen to admit more general answers when framed in the context of colored multiset Eulerian polynomials. The first question pertains to $s$-Eulerian polynomials, and the second to interpretations of $\gamma$-coefficients.

math.CO

Hyperplane Representations of Interventional Characteristic Imset Polytopes

Characteristic imsets are 0/1-vectors representing directed acyclic graphs whose edges represent direct cause-effect relations between jointly distributed random variables. A characteristic imset (CIM) polytope is the convex hull of a collection of characteristic imsets. CIM polytopes arise as feasible regions of a linear programming approach to the problem of causal disovery, which aims to infer a cause-effect structure from data. Linear optimization methods typically require a hyperplane representation of the feasible region, which has proven difficult to compute for CIM polytopes despite continued efforts. We solve this problem for CIM polytopes that are the convex hull of imsets associated to DAGs whose underlying graph of adjacencies is a tree. Our methods use the theory of toric fiber products as well as the novel notion of interventional CIM polytopes. Our solution is obtained as a corollary of a more general result for interventional CIM polytopes. The identified hyperplanes are applied to yield a linear optimization-based causal discovery algorithm for learning polytree causal networks from a combination of observational and interventional data.

math.CO

Colored Gaussian directed acyclic graphical models

We study submodels of Gaussian DAG models defined by partial homogeneity constraints imposed on the model error variances and structural coefficients. We represent these models with colored DAGs and investigate their properties for use in statistical and causal inference. Local and global Markov properties are provided and shown to characterize the colored DAG model. Additional properties relevant to causal discovery are studied, including the existence and non-existence of faithful distributions and structural identifiability. Extending prior work of Peters and B\"uhlmann and Wu and Drton, we prove structural identifiability under the assumption of homogeneous structural coefficients, as well as for a family of models with partially homogeneous structural coefficients. The latter models, termed BPEC-DAGs, capture additional causal insights by clustering the direct causes of each node into communities according to their effect on their common target. An analogue of the GES algorithm for learning BPEC-DAGs is given and evaluated on real and synthetic data. Regarding model geometry, we provide a proof of a conjecture of Sullivant which generalizes to colored DAG models, colored undirected graphical models and directed ancestral graph models. The proof yields a tool for identification of Markov properties for any rationally parameterized model with globally, rationally identifiable parameters.

math.ST

Scalable Structure Learning for Sparse Context-Specific Systems

Several approaches to graphically representing context-specific relations among jointly distributed categorical variables have been proposed, along with structure learning algorithms. While existing optimization-based methods have limited scalability due to the large number of context-specific models, the constraint-based methods are more prone to error than even constraint-based directed acyclic graph learning algorithms since more relations must be tested. We present an algorithm for learning context-specific models that scales to hundreds of variables. Scalable learning is achieved through a combination of an order-based Markov chain Monte-Carlo search and a novel, context-specific sparsity assumption that is analogous to those typically invoked for directed acyclic graphical models. Unlike previous Markov chain Monte-Carlo search methods, our Markov chain is guaranteed to have the true posterior of the variable orderings as the stationary distribution. To implement the method, we solve a first case of an open problem recently posed by Alon and Balogh. Future work solving increasingly general instances of this problem would allow our methods to learn increasingly dense models. The method is shown to perform well on synthetic data and real world examples, in terms of both accuracy and scalability.

stat.ML

Neuro-Causal Factor Analysis

Factor analysis (FA) is a statistical method for explaining how mutually dependent observed variables can be represented in terms of mutually independent latent factors, and it is widely used in the psychological, biological, and physical sciences. We revisit this classic method from the perspective of recent advances in causal structure learning and deep generative models, introducing a framework for Neuro-Causal Factor Analysis (NCFA). Our approach is fully nonparametric: it learns a directed graph between latent and observed variables, and then fits a deep generative model constrained to respect the Markov factorization of the graph. Empirically, on synthetic and real data, NCFA attains better reconstruction error compared to standard FA and better latent distribution recovery compared to a standard variational autoencoder, all with the advantages of sparser architecture, lower model complexity, and causal interpretability.

stat.ML

Triangulations of cosmological polytopes

A cosmological polytope is defined for a given Feynman diagram, and its canonical form may be used to compute the contribution of the Feynman diagram to the wavefunction of certain cosmological models. Given a subdivision of a polytope, its canonical form is obtained as a sum of the canonical forms of the facets of the subdivision. In this paper, we identify such formulas for the canonical form via algebraic techniques. It is shown that the toric ideal of every cosmological polytope admits a Gr\"obner basis with a squarefree initial ideal, yielding a regular unimodular triangulation of the polytope. In specific instances, including trees and cycles, we recover graphical characterizations of the facets of such triangulations that may be used to compute the desired canonical form. For paths and cycles, these characterizations admit simple enumeration. Hence, we obtain formulas for the normalized volume of these polytopes, extending previous observations of K\"uhne and Monin.

math.CO

Combinatorial and algebraic perspectives on the marginal independence structure of Bayesian networks

We consider the problem of estimating the marginal independence structure of a Bayesian network from observational data, learning an undirected graph we call the unconditional dependence graph. We show that unconditional dependence graphs of Bayesian networks correspond to the graphs having equal independence and intersection numbers. Using this observation, a Gr\"obner basis for a toric ideal associated to unconditional dependence graphs of Bayesian networks is given and then extended by additional binomial relations to connect the space of all such graphs. An MCMC method, called GrUES (Gr\"obner-based Unconditional Equivalence Search), is implemented based on the resulting moves and applied to synthetic Gaussian data. GrUES recovers the true marginal independence structure via a penalized maximum likelihood or MAP estimate at a higher rate than simple independence tests while also yielding an estimate of the posterior, for which the $20\%$ HPD credible sets include the true structure at a high rate for data-generating graphs with density at least $0.5$.

stat.ME

On the Edges of Characteristic Imset Polytopes

The edges of the characteristic imset polytope, $\operatorname{CIM}_p$, were recently shown to have strong connections to causal discovery as many algorithms could be interpreted as greedy restricted edge-walks, even though only a strict subset of the edges are known. To better understand the general edge structure of the polytope we describe the edge structure of faces with a clear combinatorial interpretation: for any undirected graph $G$ we have the face $\operatorname{CIM}_G$, the convex hull of the characteristic imsets of DAGs with skeleton $G$. We give a full edge-description of $\operatorname{CIM}_G$ when $G$ is a tree, leading to interesting connections to other polytopes. In particular the well-studied stable set polytope can be recovered as a face of $\operatorname{CIM}_G$ when $G$ is a tree. Building on this connection we are also able to give a description of all edges of $\operatorname{CIM}_G$ when $G$ is a cycle, suggesting possible inroads for generalization. We then introduce an algorithm for learning directed trees from data, utilizing our newly discovered edges, that outperforms classical methods on simulated Gaussian data.

math.ST

Toric Ideals of Characteristic Imsets via Quasi-Independence Gluing

Characteristic imsets are 0-1 vectors which correspond to Markov equivalence classes of directed acyclic graphs. The study of their convex hull, named the characteristic imset polytope, has led to new and interesting geometric perspectives on the important problem of causal discovery. In this paper we begin the study of the associated toric ideal. We develop a new generalization of the toric fiber product, which we call a quasi-independence gluing, and show that under certain combinatorial homogeneity conditions, one can iteratively compute a Gr\"obner basis via lifting. For faces of the characteristic imset polytope associated to trees, we apply this technique to compute a Gr\"obner basis for the associated toric ideal. We end with a study of the characteristic ideal of the cycle and propose directions for future work.

math.ST

A Transformational Characterization of Unconditionally Equivalent Bayesian Networks

We consider the problem of characterizing Bayesian networks up to unconditional equivalence, i.e., when directed acyclic graphs (DAGs) have the same set of unconditional $d$-separation statements. Each unconditional equivalence class (UEC) is uniquely represented with an undirected graph whose clique structure encodes the members of the class. Via this structure, we provide a transformational characterization of unconditional equivalence; i.e., we show that two DAGs are in the same UEC if and only if one can be transformed into the other via a finite sequence of specified moves. We also extend this characterization to the essential graphs representing the Markov equivalence classes (MECs) in the UEC. UECs partition the space of MECs and are easily estimable from marginal independence tests. Thus, a characterization of unconditional equivalence has applications in methods that involve searching the space of MECs of Bayesian networks.

stat.ML

A new characterization of discrete decomposable models

Decomposable graphical models, also known as perfect DAG models, play a fundamental role in standard approaches to probabilistic inference via graph representations in modern machine learning and statistics. However, such models are limited by the assumption that the data-generating distribution does not entail strictly context-specific conditional independence relations. The family of staged tree models generalizes DAG models so as to accommodate context-specific knowledge. We provide a new characterization of perfect discrete DAG models in terms of their staged tree representations. This characterization identifies the family of balanced staged trees as the natural generalization of discrete decomposable models to the context-specific setting.

math.ST

The Integer Decomposition Property and Weighted Projective Space Simplices

Reflexive lattice polytopes play a key role in combinatorics, algebraic geometry, physics, and other areas. One important class of lattice polytopes are lattice simplices defining weighted projective spaces. We investigate the question of when a reflexive weighted projective space simplex has the integer decomposition property. We provide a complete classification of reflexive weighted projective space simplices having the integer decomposition property for the case when there are at most three distinct non-unit weights, and conjecture a general classification for an arbitrary number of distinct non-unit weights. Further, for any weighted projective space simplex and $m\geq 1$, we define the $m$-th reflexive stabilization, a reflexive weighted projective space simplex. We prove that when $m$ is $2$ or greater, reflexive stabilizations do not have the integer decomposition property. We also prove that the Ehrhart $h^\ast$-polynomial of any sufficiently large reflexive stabilization is not unimodal and has only $1$ and $2$ as coefficients. We use this construction to generate interesting examples of reflexive weighted projective space simplices that are near the boundary of both $h^*$-unimodality and the integer decomposition property.

math.CO