arXiv ScienceSearch

arXiv subjects

Michael Fuchs

Publications and source records attributed to Michael Fuchs.

At least 19 recordsLinked to original sources

A cube-root phase transition in tree-child networks and the enumeration threshold for galled networks

We prove two surprising results about phylogenetic networks. First, we show that the structure of tree-child networks with $n$ leaves and $k$ reticulation nodes undergoes a sharp phase transition at $n^{1/3}$: if $k=o(n^{1/3})$, then a random tree-child network is almost surely a semi-simplex tree-child network, whereas if $k/n^{1/3}\rightarrow\infty$ and $k=o(n^{1/2})$, it is almost surely not. Second, we show that this result implies that the asymptotic counting formula for galled networks with $n$ leaves and a fixed number $k$ of reticulation nodes remains valid in the range $k=o(n^{1/3})$, but not beyond. This is in strong contrast to recently established results for the asymptotic counting formulas for tree-child and normal networks with $n$ leaves and $k$ reticulation nodes, which are valid in the (optimal) range $k=o(n^{1/2})$.

math.CO

Beyond Differences: Doubly Robust Meta-Learners for Ratio-Based Treatment Effects

When treatment effects are naturally expressed as ratios -- as in medicine, pricing, and marketing -- the ratio-based CATE $\tau(x) = E[Y|W=1,X=x] / E[Y|W=0,X=x]$ is the appropriate estimand. Yet existing estimators either impose a log-linear parametric structure or apply generic regression without robustness guarantees for this functional. We introduce the Q-Learner, which decomposes $\tau(x)$ into a product of two odds ratios, reducing ratio-CATE estimation for binary outcomes to two propensity classification tasks. We further derive doubly robust augmentations for both S/T- and Q-style ratio learners and characterize their distinct robustness properties. In benchmarks on seven RCT datasets, the Q-Learner is the most consistently competitive method in low-conversion regimes, where its propensity-only construction sidesteps the imbalanced regression that hurts outcome-based estimators. On four observational datasets, where propensity must be estimated and confounding cannot be ruled out, the DR learners introduced here decisively come out on top, making them practitioners' natural default for confounded observational data.

stat.ML

Combinatorial comparison of general galled trees, time-consistent galled trees, and simplex time-consistent galled trees

Rooted binary phylogenetic networks are extensions of rooted binary trees, adding reticulation nodes that are designed to represent evolutionary processes that involve hybridization events. Enumerative combinatorics studies have counted leaf-labeled phylogenetic networks in a variety of classes, finding that when the number of reticulations is fixed, the time-consistent galled trees are asymptotically less numerous than each of several network classes that had been previously examined. Here we provide enumerative results on two additional network classes: general galled trees and simplex time-consistent galled trees. We show that for a fixed number of galls, as the number of leaves goes to infinity, the asymptotic count of general galled trees is identical to that of time-consistent galled trees, whereas the count of simplex time-consistent galled trees is smaller. If the number of galls is not restricted, then the asymptotic approximations all differ: simplex time-consistent galled trees are less numerous than time-consistent galled trees, which are in turn less numerous than general galled trees. We also report a variety of additional results: recursions to count the studied networks with small numbers of leaves a fixed number of galls, as well as enumerative results for unlabeled networks in the classes that we investigate.

math.CO

The generalized Zagreb index for non-plane and plane recursive trees

The Zagreb index, which is defined as the sum of squares of degrees of the nodes of a tree, was studied in previous works by martingale techniques for random non-plane recursive trees and classes of random trees which are close to random plane recursive trees. These techniques are not easily amended to the generalized Zagreb index, which is defined similar but with squares replaced by higher powers. In this paper, we use the moment transfer approach to (i) obtain the first-order asymptotics of moments and to (ii) prove limit laws for the (suitable normalized) generalized Zagreb index for random non-plane and plane recursive trees; for the former, we show that for all higher powers the limit law is normal, for the latter, we show for cubes and fourth powers that its a non-normal law.

math.PR

A dichotomy law for certain classes of phylogenetic networks

Many classes of phylogenetic networks have been proposed in the literature. A feature of several of these classes is that if one restricts a network in the class to a subset of its leaves, then the resulting network may no longer lie within this class. This has implications for their biological applicability, since some species -- which are the leaves of an underlying evolutionary network -- may be missing (e.g., they may have become extinct, or there are no data available for them) or we may simply wish to focus attention on a subset of the species. On the other hand, certain classes of networks are `closed' when we restrict to subsets of leaves, such as (i) the classes of all phylogenetic networks or all phylogenetic trees; (ii) the classes of galled networks, simplicial networks, galled trees; and (iii) the classes of networks that have some parameter that is monotone-under-leaf-subsampling (e.g., the number of reticulations, height, etc.) bounded by some fixed value. It is easily shown that a closed subclass of phylogenetic trees is either all trees or a vanishingly small proportion of them (as the number of leaves grows). In this short paper, we explore whether this dichotomy phenomenon holds for other classes of phylogenetic networks, and their subclasses.

q-bio.PE

Enumerative combinatorics of unlabeled and labeled time-consistent galled trees

In mathematical phylogenetics, the time-consistent galled trees provide a simple class of rooted binary network structures that can be used to represent a variety of different biological phenomena. We study the enumerative combinatorics of unlabeled and labeled time-consistent galled trees. We present a new derivation via the symbolic method of the number of unlabeled time-consistent galled trees with a fixed number of leaves and a fixed number of galls. We also derive new generating functions and asymptotics for labeled time-consistent galled trees.

math.CO

Predicting the depth of the most recent common ancestor of a random sample of $k$ species: the impact of phylogenetic tree shape

We consider the following question: how close to the ancestral root of a phylogenetic tree is the most recent common ancestor of $k$ species randomly sampled from the tips of the tree? For trees having shapes predicted by the Yule-Harding model, it is known that the most recent common ancestor is likely to be close to (or equal to) the root of the full tree, even as $n$ becomes large (for $k$ fixed). However, this result does not extend to models of tree shape that more closely describe phylogenies encountered in evolutionary biology. We investigate the impact of tree shape (via the Aldous $\beta-$splitting model) to predict the number of edges that separate the most recent common ancestor of a random sample of $k$ tip species and the root of the parent tree they are sampled from. Both exact and asymptotic results are presented. We also briefly consider a variation of the process in which a random number of tip species are sampled.

q-bio.PE

The asymptotic distribution of the $k$-Robinson-Foulds dissimilarity measure on labelled trees

Motivated by applications in medical bioinformatics, Khayatian et al. (2024) introduced a family of metrics on Cayley trees (the $k$-RF distance, for $k=0, \ldots, n-2$) and explored their distribution on pairs of random Cayley trees via simulations. In this paper, we investigate this distribution mathematically, and derive exact asymptotic descriptions of the distribution of the $k$-RF metric for the extreme values $k=0$ and $k=n-2$, as $n$ becomes large. We show that a linear transform of the $0$-RF metric converges to a Poisson distribution (with mean 2) whereas a similar transform for the $(n-2)$-RF metric leads to a normal distribution (with mean $\sim ne^{-2}$). These results (together with the case $k=1$ which behaves quite differently, and $k=n-3$) shed light on the earlier simulation results, and the predictions made concerning them.

math.PR

Asymptotic enumeration of normal and hybridization networks via tree decoration

Phylogenetic networks provide a more general description of evolutionary relationships than rooted phylogenetic trees. One way to produce a phylogenetic network is to randomly place $k$ arcs between the edges of a rooted binary phylogenetic tree with $n$ leaves. The resulting directed graph may fail to be a phylogenetic network, and even when it is (and thereby a `tree-based' network), it may fail to be a tree-child or normal network. In this paper, we first show that if $k$ is fixed, the proportion of arc placements that result in a normal network tends to 1 as $n$ grows. From this result, the asymptotic enumeration of normal networks becomes straightforward and provides a transparent meaning to the combinatorial terms that arise. Moreover, the approach extends to allow $k$ to grow with $n$ (at the rate $o(n^\frac{1}{3})$), which was not handled in earlier work. We also investigate a subclass of normal networks of particular relevance in biology (hybridization networks) and establish that the same asymptotic results apply.

q-bio.PE

The $B_2$ index of galled trees

In recent years, there has been an effort to extend the classical notion of phylogenetic balance, originally defined in the context of trees, to networks. One of the most natural ways to do this is with the so-called $B_2$ index. In this paper, we study the $B_2$ index for a prominent class of phylogenetic networks: galled trees. We show that the $B_2$ index of a uniform leaf-labeled galled tree converges in distribution as the network becomes large. We characterize the corresponding limiting distribution, and provide a way to compute its moments. This is the first time that a balance index has been studied to this level of detail for a random phylogenetic network. One specificity of this work is that we use two different and independent approaches, each with its advantages: analytic combinatorics, and local limits. The analytic combinatorics approach is more direct, as it relies on standard tools; but it involves slightly more complex calculations. Because it has not previously been used to study such questions, the local limit approach requires developing an extensive framework beforehand; however, this framework is interesting in itself and can be used to tackle other similar problems.

q-bio.PE

Sackin Indices for Labeled and Unlabeled Classes of Galled Trees

The Sackin index is an important measure for the balance of phylogenetic trees. We investigate two extensions of the Sackin index to the class of galled trees and two of its subclasses (simplex galled trees and normal galled trees) where we consider both labeled and unlabeled galled trees. In all cases, we show that the mean of the Sackin index for a network which is uniformly sampled from its class is asymptotic to $\mu n^{3/2}$ for an explicit constant $\mu$. In addition, we show that the scaled Sackin index convergences weakly and with all its moments to the Airy distribution.

q-bio.PE

From Forest to Zoo: Great Ape Behavior Recognition with ChimpBehave

This paper addresses the significant challenge of recognizing behaviors in non-human primates, specifically focusing on chimpanzees. Automated behavior recognition is crucial for both conservation efforts and the advancement of behavioral research. However, it is significantly hindered by the labor-intensive process of manual video annotation. Despite the availability of large-scale animal behavior datasets, the effective application of machine learning models across varied environmental settings poses a critical challenge, primarily due to the variability in data collection contexts and the specificity of annotations. In this paper, we introduce ChimpBehave, a novel dataset featuring over 2 hours of video (approximately 193,000 video frames) of zoo-housed chimpanzees, meticulously annotated with bounding boxes and behavior labels for action recognition. ChimpBehave uniquely aligns its behavior classes with existing datasets, allowing for the study of domain adaptation and cross-dataset generalization methods between different visual settings. Furthermore, we benchmark our dataset using a state-of-the-art CNN-based action recognition model, providing the first baseline results for both within and cross-dataset settings. The dataset, models, and code can be accessed at: https://github.com/MitchFuchs/ChimpBehave

cs.CV

Galled Tree-Child Networks

We propose the class of galled tree-child networks which is obtained as intersection of the classes of galled networks and tree-child networks. For the latter two classes, (asymptotic) counting results and stochastic results have been proved with very different methods. We show that a counting result for the class of galled tree-child networks follows with similar tools as used for galled networks, however, the result has a similar pattern as the one for tree-child networks. In addition, we also consider the (suitably scaled) numbers of reticulation nodes of random galled tree-child networks and show that they are asymptotically normal distributed. This is in contrast to the limit laws of the corresponding quantities for galled networks and tree-child networks which have been both shown to be discrete.

math.CO

Counting Phylogenetic Networks with Few Reticulation Vertices: Galled and Reticulation-Visible Networks

We give exact and asymptotic counting results for the number of galled networks and reticulation-visible networks with few reticulation vertices. Our results are obtained with the component graph method, which was introduced by L. Zhang and his coauthors, and generating function techniques. For galled networks, we in addition use analytic combinatorics. Moreover, in an appendix, we consider maximally reticulated reticulation-visible networks and derive their number, too.

math.CO

The distributions under two species-tree models of the total number of ancestral configurations for matching gene trees and species trees

Given a gene-tree labeled topology $G$ and a species tree $S$, the "ancestral configurations" at an internal node $k$ of $S$ represent the combinatorially different sets of gene lineages that can be present at $k$ when all possible realizations of $G$ in $S$ are considered. Ancestral configurations have been introduced as a data structure for evaluating the conditional probability of a gene-tree labeled topology given a species tree, and their enumeration assists in describing the complexity of this computation. In the case that the gene-tree labeled topology $G=t$ matches that of the species tree $S$, by techniques of analytic combinatorics, we study distributional properties of the "total" number of ancestral configurations measured across the different nodes of a random labeled topology $t$ selected under the uniform and the Yule probability models. Under both of these probabilistic scenarios, we show that the total number $T_n$ of ancestral configurations of a random labeled topology of $n$ taxa asymptotically follows a lognormal distribution. Over uniformly distributed labeled topologies, the asymptotic growth of the mean and the variance of $T_n$ are found to satisfy $\mathbb{E}_{\rm U}[T_n] \sim 2.449 \cdot 1.333^n$ and $\mathbb{V}_{\rm U}[T_n] \sim 5.050 \cdot 1.822^n$, respectively. Under the Yule model, which assigns higher probabilities to more balanced labeled topologies, we obtain the mean $\mathbb{E}_{\rm Y}[T_n] \sim 1.425^n$ and the variance $\mathbb{V}_{\rm Y}[T_n] \sim 2.045^n$.

math.PR

Enumerative and Distributional Results for $d$-combining Tree-Child Networks

Tree-child networks are one of the most prominent network classes for modeling evolutionary processes which contain reticulation events. Several recent studies have addressed counting questions for bicombining tree-child networks in which every reticulation node has exactly two parents. We extend these studies to $d$-combining tree-child networks where every reticulation node has now $d\geq 2$ parents, and we study one-component as well as general tree-child networks. For the number of one-component networks, we derive an exact formula from which asymptotic results follow that contain a stretched exponential for $d=2$, yet not for $d \geq 3$. For general networks, we find a novel encoding by words which leads to a recurrence for their numbers. From this recurrence, we derive asymptotic results which show the appearance of a stretched exponential for all $d \geq 2$. Moreover, we also give results on the distribution of shape parameters (e.g., number of reticulation nodes, Sackin index) of a network which is drawn uniformly at random from the set of all tree-child networks with the same number of leaves. We show phase transitions depending on $d$, leading to normal, Bessel, Poisson, and degenerate distributions. Some of our results are new even in the bicombining case.

math.CO

Distribution of external branch lengths in Yule trees

The Yule branching process is a classical model for the random generation of gene tree topologies in population genetics. It generates binary ranked trees -- also called "histories" -- with a finite number $n$ of leaves. We study the lengths $\ell_1 > \ell_2 > ... > \ell_k > ...$ of the external branches of a Yule generated random history of size $n$, where the length of an external branch is defined as the rank of its parent node. When $n \rightarrow \infty$, we show that the random variable $\ell_k$, once rescaled as $\frac{n-\ell_k}{\sqrt{n/2}}$, follows a $\chi$-distribution with $2k$ degrees of freedom, with mean $\mathbb E(\ell_k) \sim n$ and variance $\mathbb V(\ell_k) \sim n \big(k-\frac{\pi k^2}{16^k} \binom{2k}{k}^2\big)$. Our results contribute to the study of the combinatorial features of Yule generated gene trees, in which external branches are associated with singleton mutations affecting individual gene copies.

math.PR

Limit Theorems for Patterns in Ranked Tree-Child Networks

We prove limit laws for the number of occurrences of a pattern on the fringe of a ranked tree-child network which is picked uniformly at random. Our results extend the limit law for cherries proved by Bienvenu et al. (2022). For patterns of height $1$ and $2$, we show that they either occur frequently (mean is asymptotically linear and limit law is normal) or sporadically (mean is asymptotically constant and limit law is Poisson) or not all (mean tends to $0$ and limit law is degenerate). We expect that these are the only possible limit laws for any fringe pattern.

math.PR