phylogenetic
434 papers tagged with this keyword
Trinets encode tree-child and level-2 phylogenetic networks
Published
• View Publication
• BIB
Phylogenetic networks generalize evolutionary trees, and are commonly used to represent evolutionary histories of species that undergo reticulate evolutionary processes such as hybridization, recombination and lateral gene transfer. Recently, there has been great interest in trying to develop methods to construct rooted phylogenetic networks from triplets, that is rooted trees on three species. However, although triplets determine or encode rooted phylogenetic trees, they do not in general encode rooted phylogenetic networks, which is a potential issue for any such method. Motivated by this fact, Huber and Moulton recently introduced trinets as a natural extension of rooted triplets to networks. In particular, they showed that level-1 phylogenetic networks are encoded by their trinets, and also conjectured that all "recoverable" rooted phylogenetic networks are encoded by their trinets. Here we prove that recoverable binary level-2 networks and binary tree-child networks are also encoded by their trinets. To do this we prove two decomposition theorems based on trinets which hold for all recoverable binary rooted phylogenetic networks. Our results provide some additional evidence in support of the conjecture that trinets encode all recoverable rooted phylogenetic networks, and could also lead to new approaches to construct phylogenetic networks from trinets.
Betti numbers of cut ideals of trees
Published in J. Alg. Stat., 4(1):108-117, 2013
• View Publication
• BIB
Cut ideals, introduced by Sturmfels and Sullivant, are used in phylogenetics and algebraic statistics. We study the minimal free resolutions of cut ideals of tree graphs. By employing basic methods from topological combinatorics, we obtain upper bounds for the Betti numbers of this type of ideals. These take the form of simple formulas on the number of vertices, which arise from the enumeration of induced subgraphs of certain incomparability graphs associated to the edge sets of trees.
Invariant polynomial functions on tensors under the action of a product of orthogonal groups
Published
• View Publication
• BIB
Let K be the product O(n_1) x O(n_2) x ... x O(n_r) of orthogonal groups. Let V the r-fold tensor product of defining representations of each orthogonal factor. We compute a stable formula for the dimension of the K-invariant algebra of degree d homogeneous polynomial functions on V. To accomplish this, we compute a formula for the number of matchings which commute with a fixed permutation. Finally, we provide formulas for the invariants and describe a bijection between a basis for the space of invariants and the isomorphism classes of certain r-regular graphs on d vertices, as well as a method of associating each invariant to other combinatorial settings such as phylogenetic trees.
Lassoing and corraling rooted phylogenetic trees
Published in Bulletin of Mathematical Biology: Volume 75, Issue 3 (2013), Page 444-465
• View Publication
• BIB
The construction of a dendogram on a set of individuals is a key component of a genomewide association study. However even with modern sequencing technologies the distances on the individuals required for the construction of such a structure may not always be reliable making it tempting to exclude them from an analysis. This, in turn, results in an input set for dendogram construction that consists of only partial distance information which raises the following fundamental question. For what subset of its leaf set can we reconstruct uniquely the dendogram from the distances that it induces on that subset. By formalizing a dendogram in terms of an edge-weighted, rooted phylogenetic tree on a pre-given finite set X with |X|>2 whose edge-weighting is equidistant and a set of partial distances on X in terms of a set L of 2-subsets of X, we investigate this problem in terms of when such a tree is lassoed, that is, uniquely determined by the elements in L. For this we consider four different formalizations of the idea of "uniquely determining" giving rise to four distinct types of lassos. We present characterizations for all of them in terms of the child-edge graphs of the interior vertices of such a tree. Our characterizations imply in particular that in case the tree in question is binary then all four types of lasso must coincide.
A simple fixed parameter tractable algorithm for computing the hybridization number of two (not necessarily binary) trees
Published
• View Publication
• BIB
Here we present a new fixed parameter tractable algorithm to compute the hybridization number r of two rooted, not necessarily binary phylogenetic trees on taxon set X in time (6^r.r!).poly(n)$, where n=|X|. The novelty of this approach is its use of terminals, which are maximal elements of a natural partial order on X, and several insights from the softwired clusters literature. This yields a surprisingly simple and practical bounded-search algorithm and offers an alternative perspective on the underlying combinatorial structure of the hybridization number problem.
Perfect taxon sampling and fixing taxon traceability: Introducing a class of phylogenetically decisive collections of taxon sets
Published
• View Publication
• BIB
Phylogenetically decisive collections of taxon sets have the property that if trees are chosen for each of their elements, as long as these trees are compatible, the resulting supertree is unique. This means that as long as the trees describing the phylogenetic relationships of the (input) species sets are compatible, they can only be combined into a common supertree in precisely one way. This setting is sometimes also referred to as \enquote{perfect taxon sampling}. While for rooted trees, the decision if a given set of input taxon sets is phylogenetically decisive can be made in polynomial time, the decision problem to determine whether a collection of taxon sets is phylogenetically decisive concerning \emph{unrooted} trees is unfortunately coNP-complete and therefore in practice hard to solve for large instances. This shows that recognizing such sets is often difficult. In this paper, we explain phylogenetic decisiveness and introduce a class of input taxon sets, namely so-called \emph{fixing taxon traceable} sets, which are guaranteed to be phylogenetically decisive and which can be recognized in polynomial time. Using both combinatorial approaches as well as simulations, we compare properties of fixing taxon traceability and phylogenetic decisiveness, e.g., by deriving lower and upper bounds for the number of quadruple sets (i.e., sets of 4-tuples) needed in the input set for each of these properties. In particular, we correct an erroneous lower bound concerning phylogenetic decisiveness from the literature.
We have implemented the algorithm to determine if a given collection of taxon sets is fixing taxon traceable in \textsf{R} and made our software package \verb+FixingTaxonTraceR+ publicly available.
Searching for Realizations of Finite Metric Spaces in Tight Spans
Published in Discrete Optimization 10 (2013), no. 4, 310-319
• View Publication
• BIB
An important problem that commonly arises in areas such as internet traffic-flow analysis, phylogenetics and electrical circuit design, is to find a representation of any given metric $D$ on a finite set by an edge-weighted graph, such that the total edge length of the graph is minimum over all such graphs. Such a graph is called an optimal realization and finding such realizations is known to be NP-hard. Recently Varone presented a heuristic greedy algorithm for computing optimal realizations. Here we present an alternative heuristic that exploits the relationship between realizations of the metric $D$ and its so-called tight span $T_D$. The tight span $T_D$ is a canonical polytopal complex that can be associated to $D$, and our approach explores parts of $T_D$ for realizations in a way that is similar to the classical simplex algorithm. We also provide computational results illustrating the performance of our approach for different types of metrics, including $l_1$-distances and two-decomposable metrics for which it is provably possible to find optimal realizations in their tight spans.
Polyhedral Combinatorics of UPGMA Cones
Published
• View Publication
• BIB
Distance-based methods such as UPGMA (Unweighted Pair Group Method with Arithmetic Mean) continue to play a significant role in phylogenetic research. We use polyhedral combinatorics to analyze the natural subdivision of the positive orthant induced by classifying the input vectors according to tree topologies returned by the algorithm. The partition lattice informs the study of UPGMA trees. We give a closed form for the extreme rays of UPGMA cones on n taxa, and compute the normalized volumes of the UPGMA cones for small n.
Keywords: phylogenetic trees, polyhedral combinatorics, partition lattice
Toric Cubes
A toric cube is a subset of the standard cube defined by binomial inequalities. These basic semialgebraic sets are precisely the images of standard cubes under monomial maps. We study toric cubes from the perspective of topological combinatorics. Explicit decompositions as CW-complexes are constructed. Their open cells are interiors of toric cubes and their boundaries are subcomplexes. The motivating example of a toric cube is the edge-product space in phylogenetics, and our work generalizes results known for that space.
On the neighbourhoods of trees
Published
• View Publication
• BIB
Tree rearrangement operations typically induce a metric on the space of phylogenetic trees. One important property of these metrics is the size of the neighbourhood, that is, the number of trees exactly one operation from a given tree. We present an expression for the size of the TBR (tree bisection and reconnection) neighbourhood, thus answering a question first posed in [Annals of Combinatorics, 5, 2001 1-15].
The maximum agreement subtree problem
Published
• View Publication
• BIB
In this paper we investigate an extremal problem on binary phylogenetic trees. Given two such trees $T_1$ and $T_2$, both with leaf-set ${1,2,...,n}$, we are interested in the size of the largest subset $S \subseteq {1,2,...,n}$ of leaves in a common subtree of $T_1$ and $T_2$. We show that any two binary phylogenetic trees have a common subtree on $Ω(\sqrt{\log{n}})$ leaves, thus improving on the previously known bound of $Ω(\log\log n)$ due to M. Steel and L. Szekely. To achieve this improved bound, we first consider two special cases of the problem: when one of the trees is balanced or a caterpillar, we show that the largest common subtree has $Ω(\log n)$ leaves. We then handle the general case by proving and applying a Ramsey-type result: that every binary tree contains either a large balanced subtree or a large caterpillar. We also show that there are constants $c, α> 0$ such that, when both trees are balanced, they have a common subtree on $c n^α$ leaves. We conjecture that it is possible to take $α= 1/2$ in the unrooted case, and both $c = 1$ and $α= 1/2$ in the rooted case.
Cycle killer... qu'est-ce que c'est? On the comparative approximability of hybridization number and directed feedback vertex set
Published
• View Publication
• BIB
We show that the problem of computing the hybridization number of two rooted binary phylogenetic trees on the same set of taxa X has a constant factor polynomial-time approximation if and only if the problem of computing a minimum-size feedback vertex set in a directed graph (DFVS) has a constant factor polynomial-time approximation. The latter problem, which asks for a minimum number of vertices to be removed from a directed graph to transform it into a directed acyclic graph, is one of the problems in Karp's seminal 1972 list of 21 NP-complete problems. However, despite considerable attention from the combinatorial optimization community it remains to this day unknown whether a constant factor polynomial-time approximation exists for DFVS. Our result thus places the (in)approximability of hybridization number in a much broader complexity context, and as a consequence we obtain that hybridization number inherits inapproximability results from the problem Vertex Cover. On the positive side, we use results from the DFVS literature to give an O(log r log log r) approximation for hybridization number, where r is the value of an optimal solution to the hybridization number problem.
Asymptotically normal distribution of some tree families relevant for phylogenetics, and of partitions without singletons
P.L. Erdos and L.A. Szekely [Adv. Appl. Math. 10(1989), 488-496] gave a bijection between rooted semilabeled trees and set partitions. L.H. Harper's results [Ann. Math. Stat. 38(1967), 410-414] on the asymptotic normality of the Stirling numbers of the second kind translates into asymptotic normality of rooted semilabeled trees with given number of vertices, when the number of internal vertices varies. The Erdos-Szekely bijection specializes to a bijection between phylogenetic trees and set partitions with classes of size \geq 2. We consider modified Stirling numbers of the second kind that enumerate partitions of a fixed set into a given number of classes of size \geq 2, and obtain their asymptotic normality as the number of classes varies. The Erdos- Szekely bijection translates this result into the asymptotic normality of the number of phylogenetic trees with given number of vertices, when the number of leaves varies. We also obtain asymptotic normality of the number of phylogenetic trees with given number of leaves and varying number of internal vertices, which make more sense to students of phylogeny. By the Erdos-Szekely bijection this means the asymptotic normality of the number of partitions of n + m elements into m classes of size \geq 2, when n is fixed and m varies. The proofs are adaptations of the techniques of L.H. Harper [ibid.]. We provide asymptotics for the relevant expectations and variances with error term O(1/n).
Optimal realisations of two-dimensional, totally-decomposable metrics
Published
• View Publication
• BIB
A realisation of a metric $d$ on a finite set $X$ is a weighted graph $(G,w)$ whose vertex set contains $X$ such that the shortest-path distance between elements of $X$ considered as vertices in $G$ is equal to $d$. Such a realisation $(G,w)$ is called optimal if the sum of its edge weights is minimal over all such realisations. Optimal realisations always exist, although it is NP-hard to compute them in general, and they have applications in areas such as phylogenetics, electrical networks and internet tomography. In [Adv. in Math. 53, 1984, 321-402] A.~Dress showed that the optimal realisations of a metric $d$ are closely related to a certain polytopal complex that can be canonically associated to $d$ called its tight-span. Moreover, he conjectured that the (weighted) graph consisting of the zero- and one-dimensional faces of the tight-span of $d$ must always contain an optimal realisation as a homeomorphic subgraph. In this paper, we prove that this conjecture does indeed hold for a certain class of metrics, namely the class of totally"=decomposable metrics whose tight-span has dimension two. As a corollary, it follows that the minimum Manhattan network problem is a special case of finding optimal realisations of two-dimensional totally-decomposable metrics.
The disentangling number for phylogenetic mixtures
Published
• View Publication
• BIB
We provide a logarithmic upper bound for the disentangling number on unordered lists of leaf labeled trees. This results is useful for analyzing phylogenetic mixture models. The proof depends on interpreting multisets of trees as high dimensional contingency tables.
On the graph labellings arising from phylogenetics
Published in Central European Journal of Mathematics, 11(9), 2013, 1577-1592
• View Publication
• BIB
We study semigroups of labellings associated to a graph. These generalize the Jukes-Cantor model and phylogenetic toric varieties defined by Buczyńska. Our main theorem bounds the degree of the generators of the semigroup by g+1 when the graph has first Betti number g. Also, we provide a series of examples where the bound is sharp.
Mathematical aspects of phylogenetic groves
Published
• View Publication
• BIB
The inference of new information on the relatedness of species by phylogenetic trees based on DNA data is one of the main challenges of modern biology. But despite all technological advances, DNA sequencing is still a time-consuming and costly process. Therefore, decision criteria would be desirable to decide a priori which data might contribute new information to the supertree which is not explicitly displayed by any input tree. A new concept, so-called groves, to identify taxon sets with the potential to construct such informative supertrees was suggested by Ané et al. in 2009. But the important conjecture that maximal groves can easily be identified in a database remained unproved and was published on the Isaac Newton Institute's list of open phylogenetic problems. In this paper, we show that the conjecture does not generally hold, but also introduce a new concept, namely 2-overlap groves, which overcomes this problem.
Affine and Projective Tree Metric Theorems
Published
• View Publication
• BIB
The tree metric theorem provides a combinatorial four point condition that characterizes dissimilarity maps derived from pairwise compatible split systems. A similar (but weaker) four point condition characterizes dissimilarity maps derived from circular split systems (Kalmanson metrics). The tree metric theorem was first discovered in the context of phylogenetics and forms the basis of many tree reconstruction algorithms, whereas Kalmanson metrics were first considered by computer scientists, and are notable in that they are a non-trivial class of metrics for which the traveling salesman problem is tractable. We present a unifying framework for these theorems based on combinatorial structures that are used for graph planarity testing. These are (projective) PC-trees, and their affine analogs, PQ-trees. In the projective case, we generalize a number of concepts from clustering theory, including hierarchies, pyramids, ultrametrics and Robinsonian matrices, and the theorems that relate them. As with tree metrics and ultrametrics, the link between PC-trees and PQ-trees is established via the Gromov product.
The Split Decomposition of a k-Dissimilarity Map
Published in Advances in Applied Mathematics, 49 (2012), Issue 1, 39-56
• View Publication
• BIB
A k-dissimilarity map on a finite set X is a function D : X \choose k \rightarrow R assigning a real value to each subset of X with cardinality k, k \geq 2. Such functions, also sometimes known as k-way dissimilarities, k-way distances, or k-semimetrics, are of interest in many areas of mathematics, computer science and classification theory, especially 2-dissimilarity maps (or distances) which are a generalisation of metrics. In this paper, we show how regular subdivisions of the kth hypersimplex can be used to obtain a canonical decomposition of a k-dissimilarity map into the sum of simpler k-dissimilarity maps arising from bipartitions or splits of X. In the special case k = 2, this is nothing other than the well-known split decomposition of a distance due to Bandelt and Dress [Adv. Math. 92 (1992), 47-105], a decomposition that is commonly to construct phylogenetic trees and networks. Furthermore, we characterise those sets of splits that may occur in the resulting decompositions of k-dissimilarity maps. As a corollary, we also give a new proof of a theorem of Pachter and Speyer [Appl. Math. Lett. 17 (2004), 615-621] for recovering k-dissimilarity maps from trees.
Non-hereditary maximum parsimony trees
Published
• View Publication
• BIB
In this paper, we investigate a conjecture by von Haeseler concerning the Maximum Parsimony method for phylogenetic estimation, which was published by the Newton Institute in Cambridge on a list of open phylogenetic problems in 2007. This conjecture deals with the question whether Maximum Parsimony trees are hereditary. The conjecture suggests that a Maximum Parsimony tree for a particular (DNA) alignment necessarily has subtrees of all possible sizes which are most parsimonious for the corresponding subalignments. We answer the conjecture affirmatively for binary alignments on five taxa but also show how to construct examples for which Maximum Parsimony trees are not hereditary. Apart from showing that a most parsimonious tree cannot generally be reduced to a most parsimonious tree on fewer taxa, we also show that compatible most parsimonious quartets do not have to provide a most parsimonious supertree. Last, we show that our results can be generalized to Maximum Likelihood for certain nucleotide substitution models.