arXiv ScienceSearch

arXiv subjects

Jeremy Sumner

Publications and source records attributed to Jeremy Sumner.

15 recordsLinked to original sources

Recombination in discrete and continuous time from the viewpoint of Markov embedding

The classic recombination equation, both in discrete and in continuous time, can be solved in a way that derives from the Markov chain of a partitioning process. Here, we revisit this structure from the point of view of the Markov embedding problem. In particular, we analyse when a discrete-time Markov matrix of recombination type can occur in a time-homogeneous Markov semigroup that is generated by a recombination rate matrix. En route, we also show that such rate matrices (or Markov generators) generally do not form a matrix algebra, but span a real Lie algebra.

math.PR

Embedding of reversible Markov matrices

The embeddability of reversible Markov matrices into time-homogeneous Markov semigroups is revisited, with some focus on simplifications and extensions. In particular, we do not demand irreducibility and consider weakly reversible matrices as well as reversible matrices with negative eigenvalues.

math.PR

On the algebra of equal-input matrices in time-inhomogeneous Markov flows

Markov matrices of equal-input type constitute a widely used model class. The corresponding equal-input generators span an interesting subalgebra of the real matrices with zero row sums. Here, we summarise some of their amazing properties and discuss the corresponding Markov embedding problem, both homogeneous and inhomogeneous in time. In particular, we derive exact and explicit solutions for time-inhomogeneous Markov flows with non-commuting generator families of equal-input type and beyond.

math.PR

Embedding of Markov matrices for $d\leqslant 4$

The embedding problem of Markov matrices in Markov semigroups is a classic problem that regained a lot of impetus and activities through recent needs in phylogeny and population genetics. Here, we give an account for dimensions $d\leqslant 4$, including a complete and simplified treatment of the case $d=3$, and derive the results in a systematic fashion, with an eye on the potential applications. Further, we reconsider the setup of the corresponding problem for time-inhomogeneous Markov chains, which is needed for real-world applications because transition rates need not be constant over time. Additional cases of this more general embedding occur for any $d\geqslant 3$. We review the known case of $d=3$ and describe the setting for future work on $d=4$.

math.PR

Evaluation of the relative performance of the subflattenings method for phylogenetic inference

The algebraic properties of flattenings and subflattenings provide direct methods for identifying edges in the true phylogeny -- and by extension the complete tree -- using pattern counts from a sequence alignment. The relatively small number of possible internal edges among a set of taxa (compared to the number of binary trees) makes these methods attractive, however more could be done to evaluate their effectiveness for inferring phylogenetic trees. This is the case particularly for subflattenings, and our work makes progress in this area. We introduce software for constructing and evaluating subflattenings for splits, utilising a number of methods to make computing subflattenings more tractable. We then present the results of simulations we have performed in order to compare the effectiveness of subflattenings to that of flattenings in terms of split score distributions, and susceptibility to possible biases. We find that subflattenings perform similarly to flattenings in terms of the distribution of split scores on the trees we examined, but may be less affected by bias arising from both split size/balance and long branch attraction. These insights are useful for developing effective algorithms to utilise these tools for the purpose of inferring phylogenetic trees.

q-bio.PE

Rearrangement Events on Circular Genomes

Early literature on genome rearrangement modelling views the problem of computing evolutionary distances as an inherently combinatorial one. In particular, attention was given to estimating distances using the minimum number of events required to transform one genome into another. In hindsight, this approach is analogous to early methods for inferring phylogenetic trees from DNA sequences such as maximum parsimony -- both are motivated by the principle that the true distance minimises evolutionary change, and both are effective if this principle is a true reflection of reality. Recent literature considers genome rearrangement under statistical models, continuing this parallel with DNA-based methods; the goal here is to use model-based methods (for example maximum likelihood techniques) to compute distance estimates that incorporate the large number of rearrangement paths that can transform one genome into another. Crucially, this approach requires one to decide upon a set of feasible rearrangement events and, in this paper, we focus on characterising well-motivated models for signed, uni-chromosomal circular genomes, where the number of regions remains fixed. Since rearrangements are often mathematically described using permutations, we isolate the sets of permutations representing rearrangements that are biologically reasonable in this context, for example inversions and translocations. We provide precise mathematical expressions for these rearrangements, and then describe them in terms of the set of cuts made in the genome when they are applied. We directly compare cuts to breakpoints, and use this concept to count the distinct rearrangement actions which apply a given number of cuts. Finally, we provide some examples of rearrangement models, and include a discussion of some questions that arise when defining plausible models.

q-bio.PE

A symmetry-inclusive algebraic approach to genome rearrangement

Of the many modern approaches to calculating evolutionary distance via models of genome rearrangement, most are tied to a particular set of genomic modelling assumptions and to a restricted class of allowed rearrangements. The "position paradigm", in which genomes are represented as permutations signifying the position (and orientation) of each region, enables a refined model-based approach, where one can select biologically plausible rearrangements and assign to them relative probabilities/costs. Here, one must further incorporate any underlying structural symmetry of the genomes into the calculations and ensure that this symmetry is reflected in the model. In our recently-introduced framework of {\em genome algebras}, each genome corresponds to an element that simultaneously incorporates all of its inherent physical symmetries. The representation theory of these algebras then provides a natural model of evolution via rearrangement as a Markov chain. Whilst the implementation of this framework to calculate distances for genomes with `practical' numbers of regions is currently computationally infeasible, we consider it to be a significant theoretical advance: one can incorporate different genomic modelling assumptions, calculate various genomic distances, and compare the results under different rearrangement models. The aim of this paper is to demonstrate some of these features.

q-bio.PE

Uniformization stable Markov models and their Jordan algebraic structure

We provide a characterisation of the continuous-time Markov models where the Markov matrices from the model can be parameterised directly in terms of the associated rate matrices (generators). That is, each Markov matrix can be expressed as the sum of the identity matrix and a rate matrix from the model. We show that the existence of an underlying Jordan algebra provides a sufficient condition, which becomes necessary for (so-called) linear models. We connect this property to the well-known uniformization procedure for continuous-time Markov chains by demonstrating that the property is equivalent to all Markov matrices from the model taking the same form as the corresponding discrete time Markov matrices in the uniformized process. We apply our results to analyse two model hierarchies practically important to phylogenetic inference, obtained by assuming (i) time-reversibility and (ii) permutation symmetry, respectively.

math.PR

A new algebraic approach to genome rearrangement models

We present a unified framework for modelling genomes and their rearrangements in a genome algebra, as elements that simultaneously incorporate all physical symmetries. Building on previous work utilising the group algebra of the symmetric group, we explicitly construct the genome algebra for the case of unsigned circular genomes with dihedral symmetry and show that the maximum likelihood estimate (MLE) of genome rearrangement distance can be validly and more efficiently performed in this setting. We then construct the genome algebra for a more general case, that is, for genomes that may be represented by elements of an arbitrary group and symmetry group, and show that the MLE computations can be performed entirely within this framework. There is no prescribed model in this framework; that is, it allows any choice of rearrangements that preserve the set of regions, along with arbitrary weights. Further, since the likelihood function is built from path probabilities -- a generalisation of path counts -- the framework may be utilised for any distance measure that is based on path probabilities.

q-bio.PE

On equal-input and monotone Markov matrices

The practically important classes of equal-input and of monotone Markov matrices are revisited, with special focus on embeddability, infinite divisibility, and mutual relations. Several uniqueness results for the classic Markov embedding problem are obtained in the process. To achieve our results, we need to employ various algebraic and geometric tools, including commutativity, permutation invariance and convexity. Of particular relevance in several demarcation results are Markov matrices that are idempotents.

math.PR

Notes on Markov embedding

The representation problem of finite-dimensional Markov matrices in Markov semigroups is revisited, with emphasis on concrete criteria for matrix subclasses of theoretical or practical relevance, such as equal-input, circulant, symmetric or doubly stochastic matrices. Here, we pay special attention to various algebraic properties of the embedding problem, and discuss the connection with the centraliser of a Markov matrix.

math.PR

Maximum likelihood estimates of pairwise rearrangement distances

Accurate estimation of evolutionary distances between taxa is important for many phylogenetic reconstruction methods. In the case of bacteria, distances can be estimated using a range of different evolutionary models, from single nucleotide polymorphisms to large-scale genome rearrangements. In the case of sequence evolution models (such as the Jukes-Cantor model and associated metric) have been used to correct pairwise distances. Similar correction methods for genome rearrangement processes are required to improve inference. Current attempts at correction fall into 3 categories: Empirical computational studies, Bayesian/MCMC approaches, and combinatorial approaches. Here we introduce a maximum likelihood estimator for the inversion distance between a pair of genomes, using the group-theoretic approach to modelling inversions introduced recently. This MLE functions as a corrected distance: in particular, we show that because of the way sequences of inversions interact with each other, it is quite possible for minimal distance and MLE distance to differently order the distances of two genomes from a third. This has obvious implications for the use of minimal distance in phylogeny reconstruction. The work also tackles the above problem allowing free rotation of the genome. Generally a frame of reference is locked, and all computation made accordingly. This work incorporates the action of the dihedral group so that distance estimates are free from any a priori frame of reference.

q-bio.PE

Is the general time-reversible model bad for molecular phylogenetics?

The general time reversible model (GTR) is presently the most popular model used in phylogentic studies. However, GTR has an undesirable mathematical property that is potentially of significant concern. It is the purpose of this article to give examples that demonstrate why this deficit may pose a problem for phylogenetic analysis and interpretation.

q-bio.PE

Lie Markov Models

Recent work has discussed the importance of multiplicative closure for the Markov models used in phylogenetics. For continuous-time Markov chains, a sufficient condition for multiplicative closure of a model class is ensured by demanding that the set of rate-matrices belonging to the model class form a Lie algebra. It is the case that some well-known Markov models do form Lie algebras and we refer to such models as "Lie Markov models". However it is also the case that some other well-known Markov models unequivocally do not form Lie algebras. In this paper, we will discuss how to generate Lie Markov models by demanding that the models have certain symmetries under nucleotide permutations. We show that the Lie Markov models include, and hence provide a unifying concept for, "group-based" and "equivariant" models. For each of two, three and four character states, the full list of Lie Markov models with maximal symmetry is presented and shown to include interesting examples that are neither group-based nor equivariant. We also argue that our scheme is pleasing in the context of applied phylogenetics, as, for a given symmetry of nucleotide substitution, it provides a natural hierarchy of models with increasing number of parameters.

q-bio.PE