arXiv++ Combinatorics

Browse math.CO papers from arXiv

phylogenetic

434 papers tagged with this keyword
2019-02-08 v2
Exchangeable and Sampling Consistent Distributions on Rooted Binary Trees
We introduce a notion of finite sampling consistency for phylogenetic trees and show that the set of finitely sampling consistent and exchangeable distributions on n leaf phylogenetic trees is a polytope. We use this polytope to show that the set of all exchangeable and infinite sampling consistent distributions on 4 leaf phylogenetic trees is exactly Aldous' beta-splitting model and give a description of some of the vertices for the polytope of distributions on 5 leaves. We also introduce a new semialgebraic set of exchangeable and sampling consistent models we call the multinomial model and use it to characterize the set of exchangeable and sampling consistent distributions.
2019-02-07 v3
Combinatorial properties of phylogenetic diversity indices
Phylogenetic diversity indices provide a formal way to apportion 'evolutionary heritage' across species. Two natural diversity indices are Fair Proportion (FP) and Equal Splits (ES). FP is also called 'evolutionary distinctiveness' and, for rooted trees, is identical to the Shapley Value (SV), which arises from cooperative game theory. In this paper, we investigate the extent to which FP and ES can differ, characterise tree shapes on which the indices are identical, and study the equivalence of FP and SV and its implications in more detail. We also define and investigate analogues of these indices on unrooted trees (where SV was originally defined), including an index that is closely related to the Pauplin representation of phylogenetic diversity.
2019-01-20
Displaying trees across two phylogenetic networks
Published in Theoretical Computer Science, 796:129-146, 2020 • View Publication • BIB
Phylogenetic networks are a generalization of phylogenetic trees to leaf-labeled directed acyclic graphs that represent ancestral relationships between species whose past includes non-tree-like events such as hybridization and horizontal gene transfer. Indeed, each phylogenetic network embeds a collection of phylogenetic trees. Referring to the collection of trees that a given phylogenetic network $N$ embeds as the display set of $N$, several questions in the context of the display set of $N$ have recently been analyzed. For example, the widely studied Tree-Containment problem asks if a given phylogenetic tree is contained in the display set of a given network. The focus of this paper are two questions that naturally arise in comparing the display sets of two phylogenetic networks. First, we analyze the problem of deciding if the display sets of two phylogenetic networks have a tree in common. Surprisingly, this problem turns out to be NP-complete even for two temporal normal networks. Second, we investigate the question of whether or not the display sets of two phylogenetic networks are equal. While we recently showed that this problem is polynomial-time solvable for a normal and a tree-child network, it is computationally hard in the general case. In establishing hardness, we show that the problem is contained in the second level of the polynomial-time hierarchy. Specifically, it is $Π_2^P$-complete. Along the way, we show that two other problems are also $Π_2^P$-complete, one of which being a generalization of Tree-Containment.
2019-01-20
Display sets of normal and tree-child networks
Published in The Electronic Journal of Combinatorics, 28, Paper 1.8, 2021 • View Publication • BIB
Phylogenetic trees canonically arise as embeddings of phylogenetic networks. We recently showed that the problem of deciding if two phylogenetic networks embed the same sets of phylogenetic trees is computationally hard, \blue{in particular, we showed it to be $Π^P_2$-complete}. In this paper, we establish a polynomial-time algorithm for this decision problem if the initial two networks consists of a normal network and a tree-child network. The running time of the algorithm is quadratic in the size of the leaf sets.
2019-01-13 v2
A class of phylogenetic networks reconstructable from ancestral profiles
Published • View Publication • BIB
Rooted phylogenetic networks provide an explicit representation of the evolutionary history of a set $X$ of sampled species. In contrast to phylogenetic trees which show only speciation events, networks can also accommodate reticulate processes (for example, hybrid evolution, endosymbiosis, and lateral gene transfer). A major goal in systematic biology is to infer evolutionary relationships, and while phylogenetic trees can be uniquely determined from various simple combinatorial data on $X$, for networks the reconstruction question is much more subtle. Here we ask when can a network be uniquely reconstructed from its `ancestral profile' (the number of paths from each ancestral vertex to each element in $X$). We show that reconstruction holds (even within the class of all networks) for a class of networks we call `orchard networks', and we provide a polynomial-time algorithm for reconstructing any orchard network from its ancestral profile. Our approach relies on establishing a structural theorem for orchard networks, which also provides for a fast (polynomial-time) algorithm to test if any given network is of orchard type. Since the class of orchard networks includes tree-sibling tree-consistent networks and tree-child networks, our result generalise reconstruction results from 2008 and 2009. Orchard networks allow for an unbounded number $k$ of reticulation vertices, in contrast to tree-sibling tree-consistent networks and tree-child networks for which $k$ is at most $2|X|-4$ and $|X|-1$, respectively.
2019-01-10 v2
The isometry group of phylogenetic tree space is $S_n$
A phylogenetic tree is an acyclic graph with distinctly labeled leaves, whose internal edges have a positive weight. Given a set of n leaves, the collection of all phylogenetic trees with this leaf set can be assembled into a metric cube complex known as phylogenetic tree space, or Billera-Holmes-Vogtmann tree space. In this largely combinatorial paper, we show that the isometry group of this space is the symmetric group on n elements. This fact is relevant to distance-based analyses of phylogenetic tree sets.
2018-12-19 v2
On Cherry-picking and Network Containment
Phylogenetic networks are used to represent evolutionary scenarios in biology and linguistics. To find the most probable scenario, it may be necessary to compare candidate networks, to distinguish different networks, and to see when one network is contained in another. In this paper, we introduce cherry-picking networks, a class of networks that can be reduced by a sequence of two graph operations. We show that some networks are uniquely determined by the sequences that reduce them---we call these the reconstructible cherry-picking networks, and further show that given two cherry-picking networks within the same reconstructible class, one is contained in the other if a sequence for the latter network reduces the former network. By restricting our scope to tree-child networks, we show that the converse of the above statement holds, thereby showing that {\sc Network Containment}, the problem of checking whether a network is contained in another, can be solved in linear time for tree-child networks. We implement this algorithm in Python and show that the linear-time theoretical bound on the input size is achievable in practice. Lastly, we provide a linear time algorithm for deciding whether two tree-child networks are isomorphic.
2018-12-17
On the Extremal Maximum Agreement Subtree Problem
Given two phylogenetic trees with the $\{1, \ldots, n\}$ leaf-set the maximum agreement subtree problem asks what is the maximum size of the subset $A \subseteq \{1, \ldots, n\}$ such that the two trees are equivalent when restricted to $A$. The long-standing extremal version of this problem focuses on the smallest number of leaves, $\mathrm{mast}(n)$, on which any two (binary and unrooted) phylogenetic trees with $n$ leaves must agree. In this work we prove that this number grows asymptotically as $Θ(\log n)$; thus closing the enduring gap between the lower and upper asymptotic bounds on $\mathrm{mast}(n)$.
Reconstructing Tree-Child Networks from Reticulate-Edge-Deleted Subnetworks
Published • View Publication • BIB
Network reconstruction lies at the heart of phylogenetic research. Two well studied classes of phylogenetic networks include tree-child networks and level-$k$ networks. In a tree-child network, every non-leaf node has a child that is a tree node or a leaf. In a level-$k$ network, the maximum number of reticulations contained in a biconnected component is $k$. Here, we show that level-$k$ tree-child networks are encoded by their reticulate-edge-deleted subnetworks, which are subnetworks obtained by deleting a single reticulation edge, if $k\geq 2$. Following this, we provide a polynomial-time algorithm for uniquely reconstructing such networks from their reticulate-edge-deleted subnetworks. Moreover, we show that this can even be done when considering subnetworks obtained by deleting one reticulation edge from each biconnected component with $k$ reticulations.
2018-11-16 v2
A tight kernel for computing the tree bisection and reconnection distance between two phylogenetic trees
Published in SIAM Journal on Discrete Mathematics, 33:1556-1574, 2019 • View Publication • BIB
In 2001 Allen and Steel showed that, if subtree and chain reduction rules have been applied to two unrooted phylogenetic trees, the reduced trees will have at most 28k taxa where k is the TBR (Tree Bisection and Reconnection) distance between the two trees. Here we reanalyse Allen and Steel's kernelization algorithm and prove that the reduced instances will in fact have at most 15k-9 taxa. Moreover we show, by describing a family of instances which have exactly 15k-9 taxa after reduction, that this new bound is tight. These instances also have no common clusters, showing that a third commonly-encountered reduction rule, the cluster reduction, cannot further reduce the size of the kernel in the worst case. To achieve these results we introduce and use "unrooted generators" which are analogues of rooted structures that have appeared earlier in the phylogenetic networks literature. Using similar argumentation we show that, for the minimum hybridization problem on two rooted trees, 9k-2 is a tight bound (when subtree and chain reduction rules have been applied) and 9k-4 is a tight bound (when, additionally, the cluster reduction has been applied) on the number of taxa, where k is the hybridization number of the two trees.
2018-11-14 v4
A structure theorem for rooted binary phylogenetic networks and its implications for tree-based networks
Published • View Publication • BIB
Attempting to recognize a tree inside a phylogenetic network is a fundamental undertaking in evolutionary analysis. In the last few years, therefore, tree-based phylogenetic networks, which are defined by a spanning tree called a subdivision tree, have attracted attention of theoretical biologists. However, the application of such networks is still not easy, due to many problems whose time complexities are not clearly understood. In this paper, we provide a general framework for solving those various old or new problems from a coherent perspective, rather than analyzing the complexity of each individual problem or developing an algorithm one by one. More precisely, we establish a structure theorem that gives a way to canonically decompose any rooted binary phylogenetic network N into maximal zig-zag trails that are uniquely determined, and use it to characterize the set of subdivision trees of N in the form of a direct product, in a way reminiscent of the structure theorem for finitely generated Abelian groups. From the main results, we derive a series of linear time and linear time delay algorithms for the following problems: given a rooted binary phylogenetic network N, 1) determine whether or not N has a subdivision tree and find one if there exists any; 2) measure the deviation of N from being tree-based; 3) compute the number of subdivision trees of N; 4) list all subdivision trees of N; and 5) find a subdivision tree to maximize or minimize a prescribed objective function. All algorithms proposed here are optimal in terms of time complexity. Our results do not only imply and unify various known results, but also answer many open questions and moreover enable novel applications, such as the estimation of a maximum likelihood tree underlying a tree-based network. The results and algorithms in this paper still hold true for a special class of rooted non-binary phylogenetic networks.
2018-10-23
Heading in the right direction? Using head moves to traverse phylogenetic network space
Published • View Publication • BIB
Head moves are a type of rearrangement moves for phylogenetic networks. They have mostly been studied as part of more encompassing types of moves, such as rSPR moves. Here, we study head moves as a type of moves on themselves. We show that the tiers ($k>0$) of phylogenetic network space are connected by local head moves. Then we show tail moves and head moves are closely related: sequences of tail moves can be converted to sequences of head moves and vice versa, changing the length by at most a constant factor. Because the tiers of network space are connected by rSPR moves, this gives a second proof of the connectivity of these tiers. Furthermore, we show that these tiers have small diameter by reproving the connectivity a third time. As the head move neighbourhood is in general small, this makes head moves a good candidate for local search heuristics. Finally we prove that finding the shortest sequence of head moves between two networks is NP-hard.
2018-10-16 v2
Lattice consensus: A partial order on phylogenetic trees that induces an associatively stable consensus method
There is a long tradition of the axiomatic study of consensus methods in phylogenetics that satisfy certain desirable properties. One recently-introduced property is associative stability, which is desirable because it confers a computational advantage, in that the consensus method only needs to be computed "pairwise". In this paper, we introduce a phylogenetic consensus method that satisfies this property, in addition to being "regular". The method is based on the introduction of a partial order on the set of rooted phylogenetic trees, itself based on the notion of a hierarchy-preserving map between trees. This partial order may be of independent interest. We call the method "lattice consensus", because it takes the unique maximal element in a lattice of trees defined by the partial order. Aside from being associatively stable, lattice consensus also satisfies the property of being Pareto on rooted triples, answering in the affirmative a question of Bryant et al (2017). We conclude the paper with an answer to another question of Bryant et al, showing that there is no regular extension stable consensus method for binary trees.
Classes of treebased networks
Published • View Publication • BIB
Recently, so-called treebased phylogenetic networks have gained considerable interest in the literature, where a treebased network is a network that can be constructed from a phylogenetic tree, called the base tree, by adding additional edges. The main aim of this manuscript is to provide some sufficient criteria for treebasedness by reducing phylogenetic networks to related graph structures. While it is generally known that deciding whether a network is treebased is NP-complete, one of these criteria, namely edgebasedness, can be verified in linear time. Surprisingly, the class of edgebased networks is closely related to a well-known family of graphs, namely the class of generalized series parallel graphs, and we will explore this relationship in full detail. Additionally, we introduce further classes of treebased networks and analyze their relationships.
Unrooted non-binary tree-based phylogenetic networks
Published • View Publication • BIB
Phylogenetic networks are a generalization of phylogenetic trees allowing for the representation of non-treelike evolutionary events such as hybridization. Typically, such networks have been analyzed based on their `level', i.e. based on the complexity of their 2-edge-connected components. However, recently the question of how `treelike' a phylogenetic network is has become the center of attention in various studies. This led to the introduction of \emph{tree-based networks}, i.e. networks that can be constructed from a phylogenetic tree, called the \emph{base tree}, by adding additional edges. While the concept of tree-basedness was originally introduced for rooted phylogenetic networks, it has recently also been considered for unrooted networks. In the present study, we compare and contrast findings obtained for unrooted \emph{binary} tree-based networks to unrooted \emph{non-binary} networks. In particular, while it is known that up to level 4 all unrooted binary networks are tree-based, we show that in the case of non-binary networks, this result only holds up to level 3.
2018-10-16 v2
A New Characterization of $\mathcal{V}$-Posets
Published • View Publication • BIB
In 2016, Hasebe and Tsujie gave a recursive characterization of the set of induced $N$-free and bowtie-free posets; Misanantenaina and Wagner studied these orders further, naming them "$\mathcal{V}$-posets". Here we offer a new characterization of $\mathcal{V}$-posets by introducing a property we refer to as autonomy. A poset $\cP$ is said to be autonomous if there exists a directed acyclic graph $D$ (with adjacency matrix $U$) whose transitive closure is $\cP$, with the property that any total ordering of the vertices of $D$ so that Gaussian elimination of $U^TU$ proceeds without row swaps is a linear extension of $\cP$. Autonomous posets arise from the theory of pressing sequences in graphs, a problem with origins in phylogenetics. The pressing sequences of a graph can be partitioned into families corresponding to posets; because of the interest in enumerating pressing sequences, we investigate when this partition has only one block, that is, when the pressing sequences are all linear extensions of a single autonomous poset. We also provide an efficient algorithm for recognition of autonomy using structural information and the forbidden subposet characterization, and we discuss a few open questions that arise in connection with these posets.
2018-08-25 v5
Ranked Schröder Trees
Published • View Publication • BIB
In biology, a phylogenetic tree is a tool to represent the evolutionary relationship between species. Unfortunately, the classical Schröder tree model is not adapted to take into account the chronology between the branching nodes. In particular, it does not answer the question: how many different phylogenetic stories lead to the creation of n species and what is the average time to get there? In this paper, we enrich this model in two distinct ways in order to obtain two ranked tree models for phylogenetics, i.e. models coding chronology. For that purpose, we first develop a model of (strongly) increasing Schröder trees, symbolically described in the classical context of increasing labeling. Then we introduce a generalization for the labeling with some unusual order constraint in Analytic Combinatorics (namely the weakly increasing trees). Although these models are direct extensions of the Schröder tree model, it appears that they are also in one-to-one correspondence with several classical combinatorial objects. Through the paper, we present these links, exhibit some parameters in typical large trees and conclude the studies with efficient uniform samplers.
2018-08-21 v2
On the uniqueness of the maximum parsimony tree for data with up to two substitutions: an extension of the classic Buneman theorem in phylogenetics
Published • View Publication • BIB
One of the main aims of phylogenetics is the reconstruction of the correct evolutionary tree when data concerning the underlying species set are given. These data typically come in the form of DNA, RNA or protein alignments, which consist of various characters (also often referred to as sites). Often, however, tree reconstruction methods based on criteria like maximum parsimony may fail to provide a unique tree for a given dataset, or, even worse, reconstruct the `wrong' tree (i.e. a tree that differs from the one that generated the data). On the other hand it has long been known that if the alignment consists of all the characters that correspond to edges of a particular tree, i.e. they all require exactly $k=1$ substitution to be realized on that tree, then this tree will be recovered by maximum parsimony methods. This is based on Buneman's theorem in mathematical phylogenetics. It is the goal of the present manuscript to extend this classic result as follows: We prove that if an alignment consists of all characters that require exactly $k=2$ substitutions on a particular tree, this tree will always be the unique maximum parsimony tree (and we also show that this can be generalized to characters which require at most $k=2$ substitutions). In particular, this also proves a conjecture based on a recently published observation by Goloboff et al. affirmatively for the special case of $k=2$.
2018-07-11
Geometric comparison of phylogenetic trees with different leaf sets
Published • View Publication • BIB
The metric space of phylogenetic trees defined by Billera, Holmes, and Vogtmann, which we refer to as BHV space, provides a natural geometric setting for describing collections of trees on the same set of taxa. However, it is sometimes necessary to analyze collections of trees on non-identical taxa sets (i.e., with different numbers of leaves), and in this context it is not evident how to apply BHV space. Davidson et al. recently approached this problem by describing a combinatorial algorithm extending tree topologies to regions in higher dimensional tree spaces, so that one can quickly compute which topologies contain a given tree as partial data. In this paper, we refine and adapt their algorithm to work for metric trees to give a full characterization of the subspace of extensions of a subtree. We describe how to apply our algorithm to define and search a space of possible supertrees and, for a collection of tree fragments with different leaf sets, to measure their compatibility.
2018-06-15 v5
The agreement distance of rooted phylogenetic networks
Published in Discrete Mathematics & Theoretical Computer Science, Vol. 21 no. 3 , Graph Theory (May 23, 2019) dmtcs:4593 • View Publication • BIB
The minimal number of rooted subtree prune and regraft (rSPR) operations needed to transform one phylogenetic tree into another one induces a metric on phylogenetic trees - the rSPR-distance. The rSPR-distance between two phylogenetic trees $T$ and $T'$ can be characterised by a maximum agreement forest; a forest with a minimum number of components that covers both $T$ and $T'$. The rSPR operation has recently been generalised to phylogenetic networks with, among others, the subnetwork prune and regraft (SNPR) operation. Here, we introduce maximum agreement graphs as an explicit representations of differences of two phylogenetic networks, thus generalising maximum agreement forests. We show that maximum agreement graphs induce a metric on phylogenetic networks - the agreement distance. While this metric does not characterise the distances induced by SNPR and other generalisations of rSPR, we prove that it still bounds these distances with constant factors.