phylogenetic
434 papers tagged with this keyword
Coalescent histories for lodgepole species trees
Published
• View Publication
• BIB
Coalescent histories are combinatorial structures that describe for a given gene tree and species tree the possible lists of branches of the species tree on which the gene tree coalescences take place. Properties of the number of coalescent histories for gene trees and species trees affect a variety of probabilistic calculations in mathematical phylogenetics. Exact and asymptotic evaluations of the number of coalescent histories, however, are known only in a limited number of cases. Here we introduce a particular family of species trees, the \emph{lodgepole} species trees $(λ_n)_{n\geq 0}$, in which tree $λ_n$ has $m=2n+1$ taxa. We determine the number of coalescent histories for the lodgepole species trees, in the case that the gene tree matches the species tree, showing that this number grows with $m!!$ in the number of taxa $m$. This computation demonstrates the existence of tree families in which the growth in the number of coalescent histories is faster than exponential. Further, it provides a substantial improvement on the lower bound for the ratio of the largest number of matching coalescent histories to the smallest number of matching coalescent histories for trees with $m$ taxa, increasing a previous bound of $(\sqrtπ / 32)[(5m-12)/(4m-6)] m \sqrt{m}$ to $[ \sqrt{m-1}/(4 \sqrt{e}) ]^{m}$. We discuss the implications of our enumerative results for phylogenetic computations.
Successful Pressing Sequences for a Bicolored Graph and Binary Matrices
Published
• View Publication
• BIB
We apply matrix theory over $\mathbb{F}_2$ to understand the nature of so-called "successful pressing sequences" of black-and-white vertex-colored graphs. These sequences arise in computational phylogenetics, where, by a celebrated result of Hannenhalli and Pevzner, the space of sortings-by-reversal of a signed permutation can be described by pressing sequences. In particular, we offer several alternative linear-algebraic and graph-theoretic characterizations of successful pressing sequences, describe the relation between such sequences, and provide bounds on the number of them. We also offer several open problems that arose as a result of the present work.
Comparing and simplifying distinct-cluster phylogenetic networks
Published in Annals of Combinatorics (2016), 1-22
• View Publication
• BIB
Phylogenetic networks are rooted acyclic directed graphs in which the leaves are identified with members of a set X of species. The cluster of a vertex is the set of leaves that are descendants of the vertex. A network is "distinct-cluster" if distinct vertices have distinct clusters. This paper focuses on the set DC(X) of distinct-cluster networks whose leaves are identified with the members of X. For a fixed X, a metric on DC(X) is defined. There is a "cluster-preserving" simplification process by which vertices or certain arcs may be removed without changing the clusters of any remaining vertices. Many of the resulting networks may be uniquely determined without regard to the order of the simplifying operations.
Facets of the Balanced Minimal Evolution Polytope
Published
• View Publication
• BIB
A phylogenetic tree is a way to organize a finite set of species, individuals or other sources of related data. The species for which we have existing DNA data make up the set of leaves of the tree. The balanced minimal evolution method of creating phylogenetic trees can be formulated as a linear programming problem, minimizing an inner product over the vertices of the BME polytope. In this paper we undertake the first steps of describing the facets of this polytope.
On the complexity of computing MP distance between binary phylogenetic trees
Published
• View Publication
• BIB
Within the field of phylogenetics there is great interest in distance measures to quantify the dissimilarity of two trees. Recently, a new distance measure has been proposed: the Maximum Parsimony (MP) distance. This is based on the difference of the parsimony scores of a single character on both trees under consideration, and the goal is to find the character which maximizes this difference. Here we show that computation of MP distance on two \emph{binary} phylogenetic trees is NP-hard. This is a highly nontrivial extension of an earlier NP-hardness proof for two multifurcating phylogenetic trees, and it is particularly relevant given the prominence of binary trees in the phylogenetics literature. As a corollary to the main hardness result we show that computation of MP distance is also hard on binary trees if the number of states available is bounded. In fact, via a different reduction we show that it is hard even if only two states are available. Finally, as a first response to this hardness we give a simple Integer Linear Program (ILP) formulation which is capable of computing the MP distance exactly for small trees (and for larger trees when only a small number of character states are available) and which is used to computationally verify several auxiliary results required by the hardness proofs.
Phylogenetic invariants for $\mathbb{Z}_3$ scheme-theoretically
Published
• View Publication
• BIB
We study phylogenetic invariants of models of evolution whose group of symmetries is the cyclic group with 3 elements. We prove that projective schemes corresponding to the ideal I of phylogenetic invariants of such a model and to its subideal I' generated by elements of degree at most 3 are the same. This is motivated by a conjecture of Sturmfels and Sullivant, which would imply that I = I'.
Combinatorics of Linked Systems of Quartet Trees
Published
• View Publication
• BIB
We apply classical quartet techniques to the problem of phylogenetic decisiveness and find a value $k$ such that all collections of at least $k$ quartets are decisive. Moreover, we prove that this bound is optimal and give a lower-bound on the probability that a collection of quartets is decisive.
Representing Partitions on Trees
Published
• View Publication
• BIB
In evolutionary biology, biologists often face the problem of constructing a phylogenetic tree on a set $X$ of species from a multiset $Π$ of partitions corresponding to various attributes of these species. One approach that is used to solve this problem is to try instead to associate a tree (or even a network) to the multiset $Σ_Π$ consisting of all those bipartitions $\{A,X-A\}$ with $A$ a part of some partition in $Π$. The rational behind this approach is that a phylogenetic tree with leaf set $X$ can be uniquely represented by the set of bipartitions of $X$ induced by its edges. Motivated by these considerations, given a multiset $Σ$ of bipartitions corresponding to a phylogenetic tree on $X$, in this paper we introduce and study the set $P(Σ)$ consisting of those multisets of partitions $Π$ of $X$ with $Σ_Π=Σ$. More specifically, we characterize when $P(Σ)$ is non-empty, and also identify some partitions in $P(Σ)$ that are of maximum and minimum size. We also show that it is NP-complete to decide when $P(Σ)$ is non-empty in case $Σ$ is an arbitrary multiset of bipartitions of $X$. Ultimately, we hope that by gaining a better understanding of the mapping that takes an arbitrary partition system $Π$ to the multiset $Σ_Π$, we will obtain new insights into the use of median networks and, more generally, split-networks to visualize sets of partitions.
On the Maximum Parsimony distance between phylogenetic trees
Published
• View Publication
• BIB
Within the field of phylogenetics there is great interest in distance measures to quantify the dissimilarity of two trees. Here, based on an idea of Bruen and Bryant, we propose and analyze a new distance measure: the Maximum Parsimony (MP) distance. This is based on the difference of the parsimony scores of a single character on both trees under consideration, and the goal is to find the character which maximizes this difference. In this article we show that this new distance is a metric and provides a lower bound to the well-known Subtree Prune and Regraft (SPR) distance. We also show that to compute the MP distance it is sufficient to consider only characters that are convex on one of the trees, and prove several additional structural properties of the distance. On the complexity side, we prove that calculating the MP distance is in general NP-hard, and identify an interesting island of tractability in which the distance can be calculated in polynomial time.
A short note on exponential-time algorithms for hybridization number
In this short note we prove that, given two (not necessarily binary) rooted phylogenetic trees T_1, T_2 on the same set of taxa X, where |X|=n, the hybridization number of T_1 and T_2 can be computed in time O^{*}(2^n) i.e. O(2^{n} poly(n)). The result also means that a Maximum Acyclic Agreement Forest (MAAF) can be computed within the same time bound.
Tropical Grassmannian and Tropical Linear Varieties from phylogenetic trees
In this paper we study tropicalization of Grassmannian and linear varieties. In particular, we study the tropical linear spaces cor- responding to the phylogenetic trees. We prove that corresponding to each subtree of the phylogenetic tree there is a point on the tropical grassmannian. We deduce a necessary and sufficient condition for it to be on the facet of the tropical linear space.
Polyhedral Covers of Tree Space
Published in SIAM Journal of Discrete Mathematics 28 (2014) 1508 - 1514
• View Publication
• BIB
The phylogenetic tree space, introduced by Billera, Holmes, and Vogtmann, is a cone over a simplicial complex. In this short article, we construct this complex from local gluings of classical polytopes, the associahedron and the permutohedron. Its homotopy is also reinterpreted and calculated based on polytope data.
Reconstructing a phylogenetic level-1 network from quartets
Published
• View Publication
• BIB
We describe a method that will reconstruct an unrooted binary phylogenetic level-1 network on n taxa from the set of all quartets containing a certain fixed taxon, in O(n^3) time. We also present a more general method which can handle more diverse quartet data, but which takes O(n^6) time. Both methods proceed by solving a certain system of linear equations over GF(2).
For a general dense quartet set (containing at least one quartet on every four taxa) our O(n^6) algorithm constructs a phylogenetic level-1 network consistent with the quartet set if such a network exists and returns an (O(n^2) sized) certificate of inconsistency otherwise. This answers a question raised by Gambette, Berry and Paul regarding the complexity of reconstructing a level-1 network from a dense quartet set.
The asymptotic enhanced negative type of finite ultrametric spaces
Published
• View Publication
• BIB
Negative type inequalities arise in the study of embedding properties of metric spaces, but they often reduce to intractable combinatorial problems. In this paper we study more quantitative versions of these inequalities involving the so-called $p$-negative type gap. In particular, we focus our attention on the class of finite ultrametric spaces which are important in areas such as phylogenetics and data mining.
Let $(X,d)$ be a given finite ultrametric space with minimum non-zero distance $α$. Then the $p$-negative type gap $Γ_{X}(p)$ of $(X,d)$ is positive for all $p \geq 0$. In this paper we compute the value of the limit \begin{eqnarray*} Γ_{X}(\infty) & = & \lim\limits_{p \rightarrow \infty} \frac{Γ_{X}(p)}{α^{p}}. \end{eqnarray*} It turns out that this value is positive and it may be given explicitly by an elegant combinatorial formula. On the basis of our calculations we are then able to characterize when $Γ_{X}(p)/ α^{p}$ is constant on $[0, \infty)$.
The determination of $Γ_{X}(\infty)$ also leads to new, asymptotically sharp, families of enhanced $p$-negative type inequalities for $(X,d)$. Indeed, suppose that $G \in (0, Γ_{X}(\infty))$. Then, for all sufficiently large $p$, we have \begin{eqnarray*} \frac{G \cdot α^{p}}{2} \left( \sum\limits_{k=1}^{n} |ζ_{k}| \right)^{2} + \sum\limits_{j,i =1}^{n} d(z_{j},z_{i})^{p} ζ_{j} ζ_{i} & \leq & 0 \end{eqnarray*} for each finite subset $\{ z_{1}, \ldots, z_{n} \} \subseteq X$ and each choice of real numbers $ζ_{1}, \ldots, ζ_{n}$ with $ζ_{1} + \cdots + ζ_{n} = 0$. We note that these results do not extend to general finite metric spaces.
A matroid associated with a phylogenetic tree
Published
• View Publication
• BIB
A (pseudo-)metric $D$ on a finite set $X$ is said to be a `tree metric' if there is a finite tree with leaf set $X$ and non-negative edge weights so that, for all $x,y \in X$, $D(x,y)$ is the path distance in the tree between $x$ and $y$. It is well known that not every metric is a tree metric. However, when some such tree exists, one can always find one whose interior edges have strictly positive edge weights and that has no vertices of degree 2, any such tree is -- up to canonical isomorphism -- uniquely determined by $D$, and one does not even need all of the distances in order to fully (re-)construct the tree's edge weights in this case. Thus, it seems of some interest to investigate which subsets of $\binom{X}{2}$ suffice to determine (`lasso') these edge weights. In this paper, we use the results of a previous paper to discuss the structure of a matroid that can be associated with an (unweighted) $X-$tree $T$ defined by the requirement that its bases are exactly the `tight edge-weight lassos' for $T$, i.e, the minimal subsets $\cl$ of $\ch$ that lasso the edge weights of $T$.
Distance-based phylogenetic methods around a polytomy
Published
• View Publication
• BIB
Distance-based phylogenetic algorithms attempt to solve the NP-hard least squares phylogeny problem by mapping an arbitrary dissimilarity map representing biological data to a tree metric. The set of all dissimilarity maps is a Euclidean space properly containing the space of all tree metrics as a polyhedral fan. Outputs of distance-based tree reconstruction algorithms such as UPGMA and Neighbor-Joining are points in the maximal cones in the fan. Tree metrics with polytomies lie at the intersections of maximal cones.
A phylogenetic algorithm divides the space of all dissimilarity maps into regions based upon which combinatorial tree is reconstructed by the algorithm. Comparison of phylogenetic methods can be done by comparing the geometry of these regions. We use polyhedral geometry to compare the local nature of the subdivisions induced by least squares phylogeny, UPGMA, and Neighbor-Joining. Our results suggest that in some circumstances, UPGMA and Neighbor-Joining poorly match least squares phylogeny when the true tree has a polytomy.
Low degree minimal generators of phylogenetic semigroups
Published
• View Publication
• BIB
The phylogenetic semigroup on a graph generalizes the Jukes-Cantor binary model on a tree. Minimal generating sets of phylogenetic semigroups have been described for trivalent trees by Buczyńska and Wiśniewski, and for trivalent graphs with first Betti number 1 by Buczyńska. We characterize degree two minimal generators of the phylogenetic semigroup on any trivalent graph. Moreover, for any graph with first Betti number 1 and for any trivalent graph with first Betti number 2 we describe the minimal generating set of its phylogenetic semigroup.
On Computing the Maximum Parsimony Score of a Phylogenetic Network
Published
• View Publication
• BIB
Phylogenetic networks are used to display the relationship of different species whose evolution is not treelike, which is the case, for instance, in the presence of hybridization events or horizontal gene transfers. Tree inference methods such as Maximum Parsimony need to be modified in order to be applicable to networks. In this paper, we discuss two different definitions of Maximum Parsimony on networks, "hardwired" and "softwired", and examine the complexity of computing them given a network topology and a character. By exploiting a link with the problem Multicut, we show that computing the hardwired parsimony score for 2-state characters is polynomial-time solvable, while for characters with more states this problem becomes NP-hard but is still approximable and fixed parameter tractable in the parsimony score. On the other hand we show that, for the softwired definition, obtaining even weak approximation guarantees is already difficult for binary characters and restricted network topologies, and fixed-parameter tractable algorithms in the parsimony score are unlikely. On the positive side we show that computing the softwired parsimony score is fixed-parameter tractable in the level of the network, a natural parameter describing how tangled reticulate activity is in the network. Finally, we show that both the hardwired and softwired parsimony score can be computed efficiently using Integer Linear Programming. The software has been made freely available.
Polyhedral computational geometry for averaging metric phylogenetic trees
Published
• View Publication
• BIB
This paper investigates the computational geometry relevant to calculations of the Frechet mean and variance for probability distributions on the phylogenetic tree space of Billera, Holmes and Vogtmann, using the theory of probability measures on spaces of nonpositive curvature developed by Sturm. We show that the combinatorics of geodesics with a specified fixed endpoint in tree space are determined by the location of the varying endpoint in a certain polyhedral subdivision of tree space. The variance function associated to a finite subset of tree space has a fixed $C^\infty$ algebraic formula within each cell of the corresponding subdivision, and is continuously differentiable in the interior of each orthant of tree space. We use this subdivision to establish two iterative methods for producing sequences that converge to the Frechet mean: one based on Sturm's Law of Large Numbers, and another based on descent algorithms for finding optima of smooth functions on convex polyhedra. We present properties and biological applications of Frechet means and extend our main results to more general globally nonpositively curved spaces composed of Euclidean orthants.
Approximation algorithms for nonbinary agreement forests
Published
• View Publication
• BIB
Given two rooted phylogenetic trees on the same set of taxa X, the Maximum Agreement Forest problem (MAF) asks to find a forest that is, in a certain sense, common to both trees and has a minimum number of components. The Maximum Acyclic Agreement Forest problem (MAAF) has the additional restriction that the components of the forest cannot have conflicting ancestral relations in the input trees. There has been considerable interest in the special cases of these problems in which the input trees are required to be binary. However, in practice, phylogenetic trees are rarely binary, due to uncertainty about the precise order of speciation events. Here, we show that the general, nonbinary version of MAF has a polynomial-time 4-approximation and a fixed-parameter tractable (exact) algorithm that runs in O(4^k poly(n)) time, where n = |X| and k is the number of components of the agreement forest minus one. Moreover, we show that a c-approximation algorithm for nonbinary MAF and a d-approximation algorithm for the classical problem Directed Feedback Vertex Set (DFVS) can be combined to yield a d(c+3)-approximation for nonbinary MAAF. The algorithms for MAF have been implemented and made publicly available.