arXiv ScienceSearch

arXiv subjects

Colby Long

Publications and source records attributed to Colby Long.

17 recordsLinked to original sources

Semialgebraic Conditions for Identifying Triangles in Phylogenetic Networks

An important consideration for a model-based method of phylogenetic network inference is the identifiability of the network parameter of the model. A recurring theme in previous works exploring this issue is that it is often difficult to identify the orientation of edges in a triangle of the network. In fact, it has been shown that for some models it is impossible to determine the orientation of triangle edges utilizing the standard algebraic technique of phylogenetic invariants. In this work, we consider one such model with a Jukes-Cantor site-substitution process and no coalescence. We give a complete semialgebraic description of three, 3-leaf Jukes-Cantor phylogenetic network models with embedded triangles. By describing these base cases, we resolve several questions about the identifiability of networks with embedded triangles. We show that for any pair of models, the intersection and set differences of the models are full-dimensional regions of the space of site-pattern probability distributions. Thus, despite being algebraically indistinguishable, these network models are not identical, nor are they identifiable (or generically identifiable). Our results also yield a straightforward biological interpretation--that the signal from a hybridization event may be immediately detectable but decays over time until it is impossible to identify the orientation of edges in the triangle of a network.

q-bio.PE

Phylogenomic Models from Tree Symmetries

A model of genomic sequence evolution on a species tree should include not only a sequence substitution process, but also a coalescent process, since different sites may evolve on different gene trees due to incomplete lineage sorting. Chifman and Kubatko initiated the study of such models, leading to the development of the SVDquartets methods of species tree inference. A key observation was that symmetries in an ultrametric species tree led to symmetries in the joint distribution of bases at the taxa. In this work, we explore the implications of such symmetry more fully, defining new models incorporating only the symmetries of this distribution, regardless of the mechanism that might have produced them. The models are thus supermodels of many standard ones with mechanistic parameterizations. We study phylogenetic invariants for the models, and establish identifiability of species tree topologies using them.

q-bio.PE

Statistical learning with phylogenetic network invariants

Phylogenetic networks provide a means of describing the evolutionary history of sets of species believed to have undergone hybridization or gene flow during their evolution. The mutation process for a set of such species can be modeled as a Markov process on a phylogenetic network. Previous work has shown that a site-pattern probability distributions from a Jukes-Cantor phylogenetic network model must satisfy certain algebraic invariants. As a corollary, aspects of the phylogenetic network are theoretically identifiable from site-pattern frequencies. In practice, because of the probabilistic nature of sequence evolution, the phylogenetic network invariants will rarely be satisfied, even for data generated under the model. Thus, using network invariants for inferring phylogenetic networks requires some means of interpreting the residuals, or deviations from zero, when observed site-pattern frequencies are substituted into the invariants. In this work, we propose a method of utilizing invariant residuals and support vector machines to infer 4-leaf level-one phylogenetic networks, from which larger networks can be reconstructed. Given data for a set of species, the support vector machine is first trained on model data to learn the patterns of residuals corresponding to different network structures to classify the network that produced the data. We demonstrate the performance of our method on simulated data from the specified model and primate data.

q-bio.PE

Distinguishing level-1 phylogenetic networks on the basis of data generated by Markov processes

Phylogenetic networks can represent evolutionary events that cannot be described by phylogenetic trees. These networks are able to incorporate reticulate evolutionary events such as hybridization, introgression, and lateral gene transfer. Recently, network-based Markov models of DNA sequence evolution have been introduced along with model-based methods for reconstructing phylogenetic networks. For these methods to be consistent, the network parameter needs to be identifiable from data generated under the model. Here, we show that the semi-directed network parameter of a triangle-free, level-1 network model with any fixed number of reticulation vertices is generically identifiable under the Jukes-Cantor, Kimura 2-parameter, or Kimura 3-parameter constraints.

q-bio.PE

Phylogenetic Networks

Phylogenetics is the study of the evolutionary relationships between organisms. One of the main challenges in the field is to take biological data for a group of organisms and to infer an evolutionary tree, a graph that represents these relationships. Developing practical and efficient methods for inferring phylogenetic trees has lead to a number of interesting mathematical questions across a variety of fields. However, due to hybridization and gene flow, a phylogenetic network may be a better representation of the evolutionary history of some groups of organisms. In this chapter, we introduce some of the basic concepts in phylogenetics, and present related undergraduate research projects on phylogenetic networks that touch on areas of graph theory and abstract algebra. In the first section, we describe several open research questions related to the combinatorics of phylogenetic networks. In the second, we describe problems related to understanding phylogenetic statistical models as algebraic varieties.

q-bio.PE

Species tree inference from genomic sequences using the log-det distance

The log-det distance between two aligned DNA sequences was introduced as a tool for statistically consistent inference of a gene tree under simple non-mixture models of sequence evolution. Here we prove that the log-det distance, coupled with a distance-based tree construction method, also permits consistent inference of species trees under mixture models appropriate to aligned genomic-scale sequences data. Data may include sites from many genetic loci, which evolved on different gene trees due to incomplete lineage sorting on an ultrametric species tree, with different time-reversible substitution processes. The simplicity and speed of distance-based inference suggests log-det based methods should serve as benchmarks for judging more elaborate and computationally-intensive species trees inference methods.

q-bio.PE

Dimensions of Group-based Phylogenetic Mixtures

In this paper we study group-based Markov models of evolution and their mixtures. In the algebreo-geometric setting, group-based phylogenetic tree models correspond to toric varieties, while their mixtures correspond to secant and join varieties. Determining properties of these secant and join varieties can aid both in model selection and establishing parameter identifiability. Here we explore the first natural geometric property of these varieties: their dimension. The expected projective dimension of the join variety of a set of varieties is one more than the sum of their dimensions. A join variety that realizes the expected dimension is nondefective. Nondefectiveness is not only interesting from a geometric point-of-view, but has been used to establish combinatorial identifiability for several classes of phylogenetic mixture models. In this paper, we focus on group-based models where the equivalence classes of identified parameters are orbits of a subgroup of the automorphism group of the group defining the model. In particular, we show that, for these group-based models, the variety corresponding to the mixture of $r$ trees with $n$ leaves is nondefective when $n \geq 2r+5$. We also give improved bounds for claw trees and give computational evidence that 2-tree and 3-tree mixtures are nondefective for small~$n$.

q-bio.PE

The effect of gene flow on coalescent-based species-tree inference

Most current methods for inferring species-level phylogenies under the coalescent model assume that no gene flow occurs following speciation. While some studies have examined the impact of gene flow on estimation accuracy for certain methods, limited analytical work has been undertaken to directly assess the potential effect of gene flow across a species phylogeny. In this paper, we consider a three-taxon isolation-with-migration model that allows gene flow between sister taxa for a brief period following speciation, as well as variation in the effective population sizes across the tree. We derive the probabilities of each of the three gene tree topologies under this model, and show that for certain choices of the gene flow and effective population size parameters, anomalous gene trees (i.e., gene trees that are discordant with the species tree but that have higher probability than the gene tree concordant with the species tree) exist. We characterize the region of parameter space producing anomalous trees, and show that the probability of the gene tree that is concordant with the species tree can be arbitrarily small. We then show that the SVDQuartets method is theoretically valid under the model of gene flow between sister taxa. We study its performance on simulated data and compare it to two other commonly-used methods for species tree inference, ASTRAL and MP-EST. The simulations show that ASTRAL and MP-EST can be statistically inconsistent when gene flow is present, while SVDQuartets performs well, though large sample sizes may be required for certain parameter choices.

q-bio.PE

Distinguishing Phylogenetic Networks

Phylogenetic networks are becoming increasingly popular in phylogenetics since they have the ability to describe a wider range of evolutionary events than their tree counterparts. In this paper, we study Markov models on phylogenetic networks and their associated geometry. We restrict our attention to large-cycle networks, networks with a single undirected cycle of length at least four. Using tools from computational algebraic geometry, we show that the semi-directed network topology is generically identifiable for Jukes-Cantor large-cycle network models.

q-bio.PE

L-infinity optimization to linear spaces and phylogenetic trees

Given a distance matrix consisting of pairwise distances between species, a distance-based phylogenetic reconstruction method returns a tree metric or equidistant tree metric (ultrametric) that best fits the data. We investigate distance-based phylogenetic reconstruction using the $l^\infty$-metric. In particular, we analyze the set of $l^\infty$-closest ultrametrics and tree metrics to an arbitrary dissimilarity map to determine its dimension and the tree topologies it represents. In the case of ultrametrics, we decompose the space of dissimilarity maps on 3 elements and on 4 elements relative to the tree topologies represented. Our approach is to first address uniqueness issues arising in $l^\infty$-optimization to linear spaces. We show that the $l^\infty$-closest point in a linear space is unique if and only if the underlying matroid of the linear space is uniform. We also give a polyhedral decomposition of $\rr^m$ based on the dimension of the set of $l^\infty$-closest points in a linear space.

math.CO

Identifiability and Reconstructibility of Species Phylogenies Under a Modified Coalescent

Coalescent models of evolution account for incomplete lineage sorting by specifying a species tree parameter which determines a distribution on gene trees. It has been shown that the unrooted topology of the species tree parameter of the multispecies coalescent is generically identifiable. Moreover, a statistically consistent reconstruction method called SVDQuartets has been developed to recover this parameter. In this paper, we describe a modified multispecies coalescent model that allows for varying effective population size and violations of the molecular clock. We show that the unrooted topology of the species tree for these models is generically identifiable and that SVDQuartets is still a statistically consistent method for inferring this parameter.

q-bio.PE

Phylogenetic trees

We introduce the package PhylogeneticTrees for Macaulay2 which allows users to compute phylogenetic invariants for group-based tree models. We provide some background information on phylogenetic algebraic geometry and show how the package PhylogeneticTrees can be used to calculate a generating set for a phylogenetic ideal as well as a lower bound for its dimension. Finally, we show how methods within the package can be used to compute a generating set for the join of any two ideals.

q-bio.PE

Initial Ideals of Pfaffian Ideals

We resolve a conjecture about a class of binomial initial ideals of $I_{2,n}$, the ideal of the Grassmannian, Gr$(2,\mathbb{C}^n$), which are associated to phylogenetic trees. For a weight vector $\omega$ in the tropical Grassmannian, $in_\omega(I_{2,n}) = J_\mathcal{T}$ is the ideal associated to the tree $\mathcal{T}$. The ideal generated by the $2r \times 2r$ subpfaffians of a generic $n \times n$ skew-symmetric matrix is precisely $I_{2,n}^{\{r-1\}}$, the $(r-1)$-secant of $I_{2,n}$. We prove necessary and sufficient conditions on the topology of $\mathcal{T}$ in order for $in_\omega(I_{2,n})^{\{2\}} = J_\mathcal{T}^{\{2\}}$. We also give a new classof prime initial ideals of the Pfaffian ideals.

math.AG

L-Infinity optimization in tropical geometry and phylogenetics

We investigate uniqueness issues that arise in $l^\infty$-optimization to linear spaces and Bergman fans of matroids. For linear spaces, we give a polyhedral decomposition of $\mathbb{R}^n$ based on the dimension of the set of $l^\infty$-nearest neighbors. This implies that the $l^\infty$-nearest neighbor in a linear space is unique if and only if the underlying matroid is uniform. For Bergman fans of matroids, we show that the set of $l^\infty$-nearest points is a tropical polytope and give an algorithm to compute its tropical vertices. A key ingredient here is a notion of topology that generalizes tree topology. These results have practical implications for distance-based phylogenetic reconstruction using the $l^\infty$-metric. We analyze the possible dimensions of the set of $l^\infty$-nearest equidistant tree metrics to an arbitrary dissimilarity map and the number of tree topologies represented in this set. For both 3 and 4-leaf trees, we decompose the space of dissimilarity maps relative to the tree topologies represented.

math.CO

Bounds on the Expected Size of the Maximum Agreement Subtree

We prove polynomial upper and lower bounds on the expected size of the maximum agreement subtree of two random binary phylogenetic trees under both the uniform distribution and Yule-Harding distribution. This positively answers a question posed in earlier work. Determining tight upper and lower bounds remains an open problem.

q-bio.PE

Tying Up Loose Strands: Defining Equations of the Strand Symmetric Model

The strand symmetric model is a phylogenetic model designed to reflect the symmetry inherent in the double-stranded structure of DNA. We show that the set of known phylogenetic invariants for the general strand symmetric model of the three leaf claw tree entirely defines the ideal. This knowledge allows one to determine the vanishing ideal of the general strand symmetric model of any trivalent tree. Our proof of the main result is computational. We use the fact that the Zariski closure of the strand symmetric model is the secant variety of a toric variety to compute the dimension of the variety. We then show that the known equations generate a prime ideal of the correct dimension using elimination theory.

q-bio.PE

Identifiability of 3-Class Jukes-Cantor Mixtures

We prove identifiability of the tree parameters of the 3-class Jukes-Cantor mixture model. The proof uses ideas from algebraic statistics, in particular: finding phylogenetic invariants that separate the varieties associated to different triples of trees; computing dimensions of the resulting phylogenetic varieties; and using the disentangling number to reduce to trees with a small number of leaves. Symbolic computation also plays a key role in handling the many different cases and finding relevant phylogenetic invariants.

q-bio.PE