phylogenetic
434 papers tagged with this keyword
From Trees to Barcodes and Back Again II: Combinatorial and Probabilistic Aspects of a Topological Inverse Problem
Published
• View Publication
• BIB
In this paper we consider two aspects of the inverse problem of how to construct merge trees realizing a given barcode. Much of our investigation exploits a recently discovered connection between the symmetric group and barcodes in general position, based on the simple observation that death order is a permutation of birth order. The first important outcome of our study is a clear combinatorial distinction between the space of phylogenetic trees (as defined by Billera, Holmes and Vogtmann) and the space of merge trees. Generic BHV trees on $n+1$ leaf nodes fall into $(2n-1)!!$ distinct strata, but the analogous number for merge trees is equal to the number of maximal chains in the lattice of partitions, i.e., $(n+1)!n!2^{-n}$. The second aspect of our study is the derivation of precise formulas for the distribution of tree realization numbers (the number of merge trees realizing a given barcode) when we assume that barcodes are sampled using a uniform distribution on the symmetric group. We are able to characterize some of the higher moments of this distribution, thanks in part to a reformulation in terms of Dirichlet convolution. This characterization provides a type of null hypothesis, apparently different from the distributions observed in real neuron data and opens the door to doing more precise science.
"Normal" phylogenetic networks may be emerging as the leading class
Published
• View Publication
• BIB
The rich and varied ways that genetic material can be passed between species has motivated extensive research into the theory of phylogenetic networks. Features that align with biological processes, or with desirable mathematical properties, have been used to define classes and prove results, with the goal of developing the theoretical foundations for network reconstruction methods. We may have now reached the point where a collection of recent results can be drawn together to make one class of network, the \emph{normal} networks, a leading contender, sitting in the sweet spot between biological relevance and mathematical tractability.
Sharp upper and lower bounds on a restricted class of convex characters
Published
• View Publication
• BIB
Let $\mathcal{T}$ be an unrooted binary tree with $n$ distinctly labelled leaves. Deriving its name from the field of phylogenetics, a convex character on $\mathcal{T}$ is simply a partition of the leaves such that the minimal spanning subtrees induced by the blocks of the partition are mutually disjoint. In earlier work Kelk and Stamoulis (Advances in Applied Mathematics 84 (2017), pp. 34--46) defined $g_k(\mathcal{T})$ as the number of convex characters where each block has at least $k$ leaves. Exact expressions were given for $g_1$ and $g_2$, where the topology of $\mathcal{T}$ turns out to be irrelevant, and it was noted that for $k \geq 3$ topological neutrality no longer holds. In this article, for every $k \geq 3$ we describe tree topologies achieving the maximum and minimum values of $g_k$ and determine corresponding expressions and exponential bounds for $g_k$. Finally, we reflect briefly on possible algorithmic applications of these results.
Non-essential arcs in phylogenetic networks
Published
• View Publication
• BIB
In the study of rooted phylogenetic networks, analyzing the set of rooted phylogenetic trees that are embedded in such a network is a recurring task. From an algorithmic viewpoint, this analysis almost always requires an exhaustive search of a particular multiset $S$ of rooted phylogenetic trees that are embedded in a rooted phylogenetic network $\mathcal{N}$. Since the size of $S$ is exponential in the number of reticulations of $\mathcal{N}$, it is consequently of interest to keep this number as small as possible but without loosing any element of $S$. In this paper, we take a first step towards this goal by introducing the notion of a non-essential arc of $\mathcal{N}$, which is an arc whose deletion from $\mathcal{N}$ results in a rooted phylogenetic network $\mathcal{N}'$ such that the sets of rooted phylogenetic trees that are embedded in $\mathcal{N}$ and $\mathcal{N}'$ are the same. We investigate the popular class of tree-child networks and characterize which arcs are non-essential. This characterization is based on a family of directed graphs. Using this novel characterization, we show that identifying and deleting all non-essential arcs in a tree-child network takes time that is cubic in the number of leaves of the network. Moreover, we show that deciding if a given arc of an arbitrary phylogenetic network is non-essential is $Π_2^P$-complete.
On the Complexity of Optimising Variants of Phylogenetic Diversity on Phylogenetic Networks
Published
• View Publication
• BIB
Phylogenetic Diversity (PD) is a prominent quantitative measure of the biodiversity of a collection of present-day species (taxa). This measure is based on the evolutionary distance among the species in the collection. Loosely speaking, if $\mathcal{T}$ is a rooted phylogenetic tree whose leaf set $X$ represents a set of species and whose edges have real-valued lengths (weights), then the PD score of a subset $S$ of $X$ is the sum of the weights of the edges of the minimal subtree of $\mathcal{T}$ connecting the species in $S$. In this paper, we define several natural variants of the PD score for a subset of taxa which are related by a known rooted phylogenetic network. Under these variants, we explore, for a positive integer $k$, the computational complexity of determining the maximum PD score over all subsets of taxa of size $k$ when the input is restricted to different classes of rooted phylogenetic networks
Combining Orthology and Xenology Data in a Common Phylogenetic Tree
Published
• View Publication
• BIB
A rooted tree $T$ with vertex labels $t(v)$ and set-valued edge labels $λ(e)$ defines maps $δ$ and $\varepsilon$ on the pairs of leaves of $T$ by setting $δ(x,y)=q$ if the last common ancestor $\text{lca}(x,y)$ of $x$ and $y$ is labeled $q$, and $m\in \varepsilon(x,y)$ if $m\inλ(e)$ for at least one edge $e$ along the path from $\text{lca}(x,y)$ to $y$. We show that a pair of maps $(δ,\varepsilon)$ derives from a tree $(T,t,λ)$ if and only if there exists a common refinement of the (unique) least-resolved vertex labeled tree $(T_δ,t_δ)$ that explains $δ$ and the (unique) least resolved edge labeled tree $(T_{\varepsilon},λ_{\varepsilon})$ that explains $\varepsilon$ (provided both trees exist). This result remains true if certain combinations of labels at incident vertices and edges are forbidden.
Phylogenetic Diversity Rankings in the Face of Extinctions: the Robustness of the Fair Proportion Index
Published
• View Publication
• BIB
Planning for the protection of species often involves difficult choices about which species to prioritize, given constrained resources. One way of prioritizing species is to consider their "evolutionary distinctiveness", i.e. their relative evolutionary isolation on a phylogenetic tree. Several evolutionary isolation metrics or phylogenetic diversity indices have been introduced in the literature, among them the so-called Fair Proportion index (also known as the "evolutionary distinctiveness" score). This index apportions the total diversity of a tree among all leaves, thereby providing a simple prioritization criterion for conservation.
Here, we focus on the prioritization order obtained from the Fair Proportion index and analyze the effects of species extinction on this ranking. More precisely, we analyze the extent to which the ranking order may change when some species go extinct and the Fair Proportion index is re-computed for the remaining taxa. We show that for each phylogenetic tree, there are edge lengths such that the extinction of one leaf per cherry completely reverses the ranking. Moreover, we show that even if only the lowest ranked species goes extinct, the ranking order may drastically change. We end by analyzing the effects of these two extinction scenarios (extinction of the lowest ranked species and extinction of one leaf per cherry) for a collection of empirical and simulated trees. In both cases, we can observe significant changes in the prioritization orders, highlighting the empirical relevance of our theoretical findings.
A Simple Linear-Time Algorithm for the Common Refinement of Rooted Phylogenetic Trees on a Common Leaf Set
Published
• View Publication
• BIB
Background. The supertree problem, i.e., the task of finding a common refinement of a set of rooted trees is an important topic in mathematical phylogenetics. The special case of a common leaf set $L$ is known to be solvable in linear time. Existing approaches refine one input tree using information of the others and then test whether the results are isomorphic.
Results. A linear-time algorithm, LinCR, for constructing the common refinement $T$ of $k$ input trees with a common leaf set is proposed that explicitly computes the parent function of $T$ in a bottom-up approach.
Conclusion. LinCR is simpler to implement than other asymptotically optimal algorithms for the problem and outperforms the alternatives in empirical comparisons.
Availability. An implementation of LinCR in Python is freely available at https://github.com/david-schaller/tralda.
Comparing the topology of phylogenetic network generators
Published
• View Publication
• BIB
Phylogenetic networks represent evolutionary history of species and can record natural reticulate evolutionary processes such as horizontal gene transfer and gene recombination. This makes phylogenetic networks a more comprehensive representation of evolutionary history compared to phylogenetic trees. Stochastic processes for generating random trees or networks are important tools in evolutionary analysis, especially in phylogeny reconstruction where they can be utilized for validation or serve as priors for Bayesian methods. However, as more network generators are developed, there is a lack of discussion or comparison for different generators. To bridge this gap, we compare a set of phylogenetic network generators by profiling topological summary statistics of the generated networks over the number of reticulations and comparing the topological profiles.
Measuring tree balance using symmetry nodes -- a new balance index and its extremal properties
Published
• View Publication
• BIB
Effects like selection in evolution as well as fertility inheritance in the development of populations can lead to a higher degree of asymmetry in evolutionary trees than expected under a null hypothesis. To identify and quantify such influences, various balance indices were proposed in the phylogenetic literature and have been in use for decades.
However, so far no balance index was based on the number of \emph{symmetry nodes}, even though symmetry nodes play an important role in other areas of mathematical phylogenetics and despite the fact that symmetry nodes are a quite natural way to measure balance or symmetry of a given tree.
The aim of this manuscript is thus twofold: First, we will introduce the \emph{symmetry nodes index} as an index for measuring balance of phylogenetic trees and analyze its extremal properties. We also show that this index can be calculated in linear time. This new index turns out to be a generalization of a simple and well-known balance index, namely the \emph{cherry index}, as well as a specialization of another, less established, balance index, namely \emph{Rogers' $J$ index}. Thus, it is the second objective of the present manuscript to compare the new symmetry nodes index to these two indices and to underline its advantages. In order to do so, we will derive some extremal properties of the cherry index and Rogers' $J$ index along the way and thus complement existing studies on these indices. Moreover, we used the programming language \textsf{R} to implement all three indices in the software package \textsf{symmeTree}, which has been made publicly available.
Compatibility of Partitions with Trees, Hierarchies, and Split Systems
Published
• View Publication
• BIB
The question whether a partition $\mathcal{P}$ and a hierarchy $\mathcal{H}$ or a tree-like split system $\mathfrak{S}$ are compatible naturally arises in a wide range of classification problems. In the setting of phylogenetic trees, one asks whether the sets of $\mathcal{P}$coincide with leaf sets of connected components obtained by deleting some edges from the tree $T$ that represents $\mathcal{H}$ or $\mathfrak{S}$, respectively. More generally, we ask whether a refinement $T^*$ of $T$ exists such that $T^*$ and $\mathcal{P}$ are compatible in this sense. The latter is closely related to the question as to whether there exists a tree at all that is compatible with $\mathcal{P}$. We report several characterizations for (refinements of) hierarchies and split systems that are compatible with (systems of) partitions. In addition, we provide a linear-time algorithm to check whether refinements of trees and a given partition are compatible. The latter problem becomes NP-complete but fixed-parameter tractable if a system of partitions is considered instead of a single partition. In this context, we also explore the close relationship of the concept of compatibility and so-called Fitch maps.
Distinguishing Level-2 Phylogenetic Networks Using Phylogenetic Invariants
In phylogenetics, it is important for the phylogenetic network model parameters to be identifiable so that the evolutionary histories of a group of species can be consistently inferred. However, as the complexity of the phylogenetic network models grows, the identifiability of network models becomes increasingly difficult to analyze. As an attempt to analyze the identifiability of network models, we check whether two networks are distinguishable. In this paper, we specifically study the distinguishability of phylogenetic network models associated with level-2 networks. Using an algebraic approach, namely using discrete Fourier transformation, we present some results on the distinguishability of some level-2 networks, which generalize earlier work on the distinguishability of level-1 networks. In particular, we study simple and semisimple level-2 networks. Simple and semisimple level-2 networks can be thought as generalizations of level-1 sunlet and cycle networks, respectively. Moreover, we also compare the varieties associated with semisimple level-2 and cycle networks.
Tree Topologies along a Tropical Line Segment
Published
• View Publication
• BIB
Tropical geometry with the max-plus algebra has been applied to statistical learning models over tree spaces because geometry with the tropical metric over tree spaces has some nice properties such as convexity in terms of the tropical metric. One of the challenges in applications of tropical geometry to tree spaces is the difficulty interpreting outcomes of statistical models with the tropical metric. This paper focuses on combinatorics of tree topologies along a tropical line segment, an intrinsic geodesic with the tropical metric, between two phylogenetic trees over the tree space and we show some properties of a tropical line segment between two trees. Specifically we show that a probability of a tropical line segment of two randomly chosen trees going through the origin (the star tree) is zero if the number of leave is greater than four, and we also show that if two given trees differ only one nearest neighbor interchange (NNI) move, then the tree topology of a tree in the tropical line segment between them is the same tree topology of one of these given two trees with possible zero branch lengths.
Counting Phylogenetic Networks with Few Reticulation Vertices: A Second Approach
Published
• View Publication
• BIB
Tree-child networks, one of the prominent network classes in phylogenetics, have been introduced for the purpose of modeling reticulate evolution. Recently, the first author together with Gittenberger and Mansouri (2019) showed that the number ${\rm TC}_{\ell,k}$ of tree-child networks with $\ell$ leaves and $k$ reticulation vertices has the first-order asymptotics \[ {\rm TC}_{\ell,k}\sim c_k\left(\frac{2}{e}\right)^{\ell}\ell^{\ell+2k-1},\qquad (\ell\rightarrow\infty). \] Moreover, they also computed $c_k$ for $k=1,2,$ and $3$. In this short note, we give a second approach to the above result which is based on a recent (algorithmic) approach for the counting of tree-child networks due to Cardona and Zhang (2020). This second approach is also capable of giving a simple, closed-form expression for $c_k$, namely, $c_k=2^{k-1}\sqrt{2}/k!$ for all $k\geq 0$.
Arc-Completion of 2-Colored Best Match Graphs to Binary-Explainable Best Match Graphs
Published
• View Publication
• BIB
Best match graphs (BMGs) are vertex-colored digraphs that naturally arise in mathematical phylogenetics to formalize the notion of evolutionary closest genes w.r.t. an a priori unknown phylogenetic tree. BMGs are explained by unique least resolved trees. We prove that the property of a rooted, leaf-colored tree to be least resolved for some BMG is preserved by the contraction of inner edges. For the special case of two-colored BMGs, this leads to a characterization of the least resolved trees (LRTs) of binary-explainable trees and a simple, polynomial-time algorithm for the minimum cardinality completion of the arc set of a BMG to reach a BMG that can be explained by a binary tree.
From Modular Decomposition Trees to Rooted Median Graphs
Published
• View Publication
• BIB
The modular decomposition of a symmetric map $δ\colon X\times X \to Υ$ (or, equivalently, a set of symmetric binary relations, a 2-structure, or an edge-colored undirected graph) is a natural construction to capture key features of $δ$ in labeled trees. A map $δ$ is explained by a vertex-labeled rooted tree $(T,t)$ if the label $δ(x,y)$ coincides with the label of the last common ancestor of $x$ and $y$ in $T$, i.e., if $δ(x,y)=t(\mathrm{lca}(x,y))$. Only maps whose modular decomposition does not contain prime nodes, i.e., the symbolic ultrametrics, can be exaplained in this manner. Here we consider rooted median graphs as a generalization to (modular decomposition) trees to explain symmetric maps. We first show that every symmetric map can be explained by "extended" hypercubes and half-grids. We then derive a a linear-time algorithm that stepwisely resolves prime vertices in the modular decomposition tree to obtain a rooted and labeled median graph that explains a given symmetric map $δ$. We argue that the resulting "tree-like" median graphs may be of use in phylogenetics as a model of evolutionary relationships.
Heuristic Algorithms for Best Match Graph Editing
Published
• View Publication
• BIB
Best match graphs (BMGs) are a class of colored digraphs that naturally appear in mathematical phylogenetics and can be approximated with the help of similarity measures between gene sequences, albeit not without errors. The corresponding graph editing problem can be used as a means of error correction. Since the arc set modification problems for BMGs are NP-complete, efficient heuristics are needed if BMGs are to be used for the practical analysis of biological sequence data. Since BMGs have a characterization in terms of consistency of a certain set of rooted triples, we consider heuristics that operate on triple sets. As an alternative, we show that there is a close connection to a set partitioning problem that leads to a class of top-down recursive algorithms that are similar to Aho's supertree algorithm and give rise to BMG editing algorithms that are consistent in the sense that they leave BMGs invariant. Extensive benchmarking shows that community detection algorithms for the partitioning steps perform best for BMG editing.
A branching process approach to level-$k$ phylogenetic networks
Published
• View Publication
• BIB
The mathematical analysis of random phylogenetic networks via analytic and algorithmic methods has received increasing attention in the past years. In the present work we introduce branching process methods to their study. This approach appears to be new in this context. Our main results focus on random level-$k$ networks with $n$ labelled leaves. Although the number of reticulation vertices in such networks is typically linear in $n$, we prove that their asymptotic global and local shape is tree-like in a well-defined sense. We show that the depth process of vertices in a large network converges towards a Brownian excursion after rescaling by $n^{-1/2}$. We also establish Benjamini--Schramm convergence of large random level-$k$ networks towards a novel random infinite network.
A survey of the monotonicity and non-contradiction of consensus methods and supertree methods
Published
• View Publication
• BIB
In a recent study, Bryant, Francis and Steel investigated the concept of \enquote{future-proofing} consensus methods in phylogenetics. That is, they investigated if such methods can be robust against the introduction of additional data like added trees or new species. In the present manuscript, we analyze consensus methods under a different aspect of introducing new data, namely concerning the discovery of new clades. In evolutionary biology, often formerly unresolved clades get resolved by refined reconstruction methods or new genetic data analyses. In our manuscript we investigate which properties of consensus methods can guarantee that such new insights do not disagree with previously found consensus trees, but merely refine them, a property termed \emph{monotonicity}. Along the lines of analyzing monotonicity, we also study two {established} supertree methods, namely Matrix Representation with Parsimony (MRP) and Matrix Representation with Compatibility (MRC), which have also been suggested as consensus methods in the literature. While we (just like Bryant, Francis and Steel in their recent study) unfortunately have to conclude some negative answers concerning general consensus methods, we also state some relevant and positive results concerning the majority rule ($\mathtt{MR}$) and strict consensus methods, which are amongst the most frequently used consensus methods. Moreover, we show that there exist infinitely many consensus methods which are monotonic and have some other desirable properties.
\textbf{Keywords:} consensus tree, phylogenetics, majority rule, tree refinement, matrix representation with parsimony
\textbf{MSC:} C92B05, 05C05
Level-$2$ networks from shortest and longest distances
Published
• View Publication
• BIB
Recently it was shown that a certain class of phylogenetic networks, called level-$2$ networks, cannot be reconstructed from their associated distance matrices. In this paper, we show that they can be reconstructed from their induced shortest and longest distance matrices. That is, if two level-$2$ networks induce the same shortest and longest distance matrices, then they must be isomorphic. We further show that level-$2$ networks are reconstructible from their shortest distance matrices if and only if they do not contain a subgraph from a family of graphs. A generator of a network is the graph obtained by deleting all pendant subtrees and suppressing degree-$2$ vertices. We also show that networks with a leaf on every generator side is reconstructible from their induced shortest distance matrix, regardless of level.