phylogenetic
434 papers tagged with this keyword
Slim Sets of Binary Trees
A classical problem in phylogenetic tree analysis is to decide whether there is a phylogenetic tree $T$ that contains all information of a given collection $\cP$ of phylogenetic trees. If the answer is "yes" we say that $\cP$ is compatible and $T$ displays $\cP$. This decision problem is NP-complete even if all input trees are quartets, that is binary trees with exactly four leaves. In this paper, we prove a sufficient condition for a set of binary phylogenetic trees to be compatible. That result is used to give a short and self-contained proof of the known characterization of quartet sets of minimal cardinality which are displayed by a unique phylogenetic tree.
Restricted trees: simplifying networks with bottlenecks
Published in Bulletin of Mathematical Biology (2011) 73, 2322-2338
• View Publication
• BIB
Suppose N is a phylogenetic network indicating a complicated relationship among individuals and taxa. Often of interest is a much simpler network, for example, a species tree T, that summarizes the most fundamental relationships. The meaning of a species tree is made more complicated by the recent discovery of the importance of hybridizations and lateral gene transfers. Hence it is desirable to describe uniform well-defined procedures that yield a tree given a network N. A useful tool toward this end is a connected surjective digraph (CSD) map f from N to N' where N' is generally a much simpler network than N. A set W of vertices in N is "restricted" if there is at most one vertex from which there is an arc into W, thus yielding a bottleneck in N. A CSD map f from N to N' is "restricted" if the inverse image of each vertex in N' is restricted in N. This paper describes a uniform procedure that, given a network N, yields a well-defined tree called the "restricted tree" of N. There is a restricted CSD map from N to the restricted tree. Many relationships in the tree can be proved to appear also in N.
CSD Homomorphisms Between Phylogenetic Networks
Published in IEEE/ACM Transactions on Computational Biology and Bioinformatics (2012) 9: 1128-1138
• View Publication
• BIB
Since Darwin, species trees have been used as a simplified description of the relationships which summarize the complicated network $N$ of reality. Recent evidence of hybridization and lateral gene transfer, however, suggest that there are situations where trees are inadequate. Consequently it is important to determine properties that characterize networks closely related to $N$ and possibly more complicated than trees but lacking the full complexity of $N$.
A connected surjective digraph map (CSD) is a map $f$ from one network $N$ to another network $M$ such that every arc is either collapsed to a single vertex or is taken to an arc, such that $f$ is surjective, and such that the inverse image of a vertex is always connected. CSD maps are shown to behave well under composition. It is proved that if there is a CSD map from $N$ to $M$, then there is a way to lift an undirected version of $M$ into $N$, often with added resolution. A CSD map from $N$ to $M$ puts strong constraints on $N$.
In general, it may be useful to study classes of networks such that, for any $N$, there exists a CSD map from $N$ to some standard member of that class.
Optimality of the Neighbor Joining Algorithm and Faces of the Balanced Minimum Evolution Polytope
Published
• View Publication
• BIB
Balanced minimum evolution (BME) is a statistically consistent distance-based method to reconstruct a phylogenetic tree from an alignment of molecular data. In 2000, Pauplin showed that the BME method is equivalent to optimizing a linear functional over the BME polytope, the convex hull of the BME vectors obtained from Pauplin's formula applied to all binary trees. The BME method is related to the Neighbor Joining (NJ) algorithm, now known to be a greedy optimization of the BME principle. Further, the NJ and BME algorithms have been studied previously to understand when the NJ Algorithm returns a BME tree for small numbers of taxa. In this paper we aim to elucidate the structure of the BME polytope and strengthen knowledge of the connection between the BME method and NJ Algorithm. We first prove that any subtree-prune-regraft move from a binary tree to another binary tree corresponds to an edge of the BME polytope. Moreover, we describe an entire family of faces parametrized by disjoint clades. We show that these {\em clade-faces} are smaller dimensional BME polytopes themselves. Finally, we show that for any order of joining nodes to form a tree, there exists an associated distance matrix (i.e., dissimilarity map) for which the NJ Algorithm returns the BME tree. More strongly, we show that the BME cone and every NJ cone associated to a tree $T$ have an intersection of positive measure.
Polyhedral geometry of Phylogenetic Rogue Taxa
Published
• View Publication
• BIB
It is well known among phylogeneticists that adding an extra taxon (e.g. species) to a data set can alter the structure of the optimal phylogenetic tree in surprising ways. However, little is known about this "rogue taxon" effect. In this paper we characterize the behavior of balanced minimum evolution (BME) phylogenetics on data sets of this type using tools from polyhedral geometry. First we show that for any distance matrix there exist distances to a "rogue taxon" such that the BME-optimal tree for the data set with the new taxon does not contain any nontrivial splits (bipartitions) of the optimal tree for the original data. Second, we prove a theorem which restricts the topology of BME-optimal trees for data sets of this type, thus showing that a rogue taxon cannot have an arbitrary effect on the optimal tree. Third, we construct polyhedral cones computationally which give complete answers for BME rogue taxon behavior when our original data fits a tree on four, five, and six taxa. We use these cones to derive sufficient conditions for rogue taxon behavior for four taxa, and to understand the frequency of the rogue taxon effect via simulation.
The Geometry of the Neighbor-Joining Algorithm for Small Trees
Published in It is published in the proceedings of Algebraic Biology, Springer LNC Series (2008), p82-96
• View Publication
• BIB
In 2007, Eickmeyer et al. showed that the tree topologies outputted by the Neighbor-Joining (NJ) algorithm and the balanced minimum evolution (BME) method for phylogenetic reconstruction are each determined by a polyhedral subdivision of the space of dissimilarity maps ${\R}^{n \choose 2}$, where $n$ is the number of taxa. In this paper, we will analyze the behavior of the Neighbor-Joining algorithm on five and six taxa and study the geometry and combinatorics of the polyhedral subdivision of the space of dissimilarity maps for six taxa as well as hyperplane representations of each polyhedral subdivision. We also study simulations for one of the questions stated by Eickmeyer et al., that is, the robustness of the NJ algorithm to small perturbations of tree metrics, with tree models which are known to be hard to be reconstructed via the NJ algorithm.
A Note on Encodings of Phylogenetic Networks of Bounded Level
Published
• View Publication
• BIB
Driven by the need for better models that allow one to shed light into the question how life's diversity has evolved, phylogenetic networks have now joined phylogenetic trees in the center of phylogenetics research. Like phylogenetic trees, such networks canonically induce collections of phylogenetic trees, clusters, and triplets, respectively. Thus it is not surprising that many network approaches aim to reconstruct a phylogenetic network from such collections. Related to the well-studied perfect phylogeny problem, the following question is of fundamental importance in this context: When does one of the above collections encode (i.e. uniquely describe) the network that induces it? In this note, we present a complete answer to this question for the special case of a level-1 (phylogenetic) network by characterizing those level-1 networks for which an encoding in terms of one (or equivalently all) of the above collections exists. Given that this type of network forms the first layer of the rich hierarchy of level-k networks, k a non-negative integer, it is natural to wonder whether our arguments could be extended to members of that hierarchy for higher values for k. By giving examples, we show that this is not the case.
Computing Geodesic Distances in Tree Space
Published
• View Publication
• BIB
We present two algorithms for computing the geodesic distance between phylogenetic trees in tree space, as introduced by Billera, Holmes, and Vogtmann (2001). We show that the possible combinatorial types of shortest paths between two trees can be compactly represented by a partially ordered set. We calculate the shortest distance along each candidate path by converting the problem into one of finding the shortest path through a certain region of Euclidean space. In particular, we show there is a linear time algorithm for finding the shortest path between a point in the all positive orthant and a point in the all negative orthant of R^k contained in the subspace of R^k consisting of all orthants with the first i coordinates non-positive and the remaining coordinates non-negative for 0 <= i <= k.
Boxicity of Leaf Powers
Published
• View Publication
• BIB
The boxicity of a graph G, denoted as box(G) is defined as the minimum integer t such that G is an intersection graph of axis-parallel t-dimensional boxes. A graph G is a k-leaf power if there exists a tree T such that the leaves of the tree correspond to the vertices of G and two vertices in G are adjacent if and only if their corresponding leaves in T are at a distance of at most k. Leaf powers are a subclass of strongly chordal graphs and are used in the construction of phylogenetic trees in evolutionary biology. We show that for a k-leaf power G, box(G)\leq k-1. We also show the tightness of this bound by constructing a k-leaf power with boxicity equal to k-1. This result implies that there exists strongly chordal graphs with arbitrarily high boxicity which is somewhat counterintuitive.
Isomorphism and Symmetries in Random Phylogenetic Trees
The probability that two randomly selected phylogenetic trees of the same size are isomorphic is found to be asymptotic to a decreasing exponential modulated by a polynomial factor. The number of symmetrical nodes in a random phylogenetic tree of large size obeys a limiting Gaussian distribution, in the sense of both central and local limits. The probability that two random phylogenetic trees have the same number of symmetries asymptotically obeys an inverse square-root law. Precise estimates for these problems are obtained by methods of analytic combinatorics, involving bivariate generating functions, singularity analysis, and quasi-powers approximations.
Least Squares Methods for Equidistant Tree Reconstruction
UPGMA is a heuristic method identifying the least squares equidistant phylogenetic tree given empirical distance data among $n$ taxa. We study this classic algorithm using the geometry of the space of all equidistant trees with $n$ leaves, also known as the Bergman complex of the graphical matroid for the complete graph $K_n$. We show that UPGMA performs an orthogonal projection of the data onto a maximal cell of the Bergman complex. We also show that the equidistant tree with the least (Euclidean) distance from the data is obtained from such an orthogonal projection, but not necessarily given by UPGMA. Using this geometric information we give an extension of the UPGMA algorithm. We also present a branch and bound method for finding the best equidistant tree. Finally, we prove that there are distance data among $n$ taxa which project to at least $(n-1)!$ equidistant trees.
Sagbi Bases of Cox-Nagata Rings
Published
• View Publication
• BIB
We degenerate Cox-Nagata rings to toric algebras by means of sagbi bases induced by configurations over the rational function field. For del Pezzo surfaces, this degeneration implies the Batyrev-Popov conjecture that these rings are presented by ideals of quadrics. For the blow-up of projective n-space at n+3 points, sagbi bases of Cox-Nagata rings establish a link between the Verlinde formula and phylogenetic algebraic geometry, and we use this to answer questions due to D'Cruz-Iarobbino and Buczynska-Wisniewski. Inspired by the zonotopal algebras of Holtz and Ron, our study emphasizes explicit computations, and offers a new approach to Hilbert functions of fat points.
Combinatorics of least squares trees
Published
• View Publication
• BIB
A recurring theme in the least squares approach to phylogenetics has been the discovery of elegant combinatorial formulas for the least squares estimates of edge lengths. These formulas have proved useful for the development of efficient algorithms, and have also been important for understanding connections among popular phylogeny algorithms. For example, the selection criterion of the neighbor-joining algorithm is now understood in terms of the combinatorial formulas of Pauplin for estimating tree length.
We highlight a phylogenetically desirable property that weighted least squares methods should satisfy, and provide a complete characterization of methods that satisfy the property. The necessary and sufficient condition is a multiplicative four point condition that the the variance matrix needs to satisfy. The proof is based on the observation that the Lagrange multipliers in the proof of the Gauss--Markov theorem are tree-additive. Our results generalize and complete previous work on ordinary least squares, balanced minimum evolution and the taxon weighted variance model. They also provide a time optimal algorithm for computation.
The space of tropically collinear points is shellable
Published in Collectanea Mathematica 60, 1 (2009), pp 63-77
• View Publication
• BIB
The space T_{d,n} of n tropically collinear points in a fixed tropical projective space TP^{d-1} is equivalent to the tropicalization of the determinantal variety of matrices of rank at most 2, which consists of real d x n matrices of tropical or Kapranov rank at most 2, modulo projective equivalence of columns. We show that it is equal to the image of the moduli space M_{0,n}(TP^{d-1},1) of n-marked tropical lines in TP^{d-1} under the evaluation map. Thus we derive a natural simplicial fan structure for T_{d,n} using a simplicial fan structure of M_{0,n}(TP^{d-1},1) which coincides with that of the space of phylogenetic trees on d+n taxa. The space of phylogenetic trees has been shown to be shellable by Trappmann and Ziegler. Using a similar method, we show that T_{d,n} is shellable with our simplicial fan structure and compute the homology of the link of the origin. The shellability of T_{d,n} has been conjectured by Develin in 2005.
Phylogenetic networks form partial trees
A contemporary and fundamental problem faced by many evolutionary biologists is how to puzzle together a collection $\mathcal P$ of partial trees (leaf-labelled trees whose leaves are bijectively labelled by species or, more generally, taxa, each supported by e. g. a gene) into an overall parental structure that displays all trees in $\mathcal P$. This already difficult problem is complicated by the fact that the trees in $\mathcal P$ regularly support conflicting phylogenetic relationships and are not on the same but only overlapping taxa sets. A desirable requirement on the sought after parental structure therefore is that it can accommodate the observed conflicts. Phylogenetic networks are a popular tool capable of doing precisely this. However, not much is known about how to construct such networks from partial trees, a notable exception being the $Z$-closure super-network approach and the recently introduced $Q$-imputation approach. Here, we propose the usage of closure rules to obtain such a network. In particular, we introduce the novel $Y$-closure rule and show that this rule on its own or in combination with one of Meacham's closure rules (which we call the $M$-rule) has some very desirable theoretical properties. In addition, we use the $M$- and $Y$-rule to explore the dependency of Rivera et al.'s ``ring of life'' on the fact that the underpinning phylogenetic trees are all on the same data set. Our analysis culminates in the presentation of a collection of induced subtrees from which this ring can be reconstructed.
Affine Buildings and Tropical Convexity
Published in Albanian J. Math. 1 (2007), no. 4, 187--211
• View Publication
• BIB
The notion of convexity in tropical geometry is closely related to notions of convexity in the theory of affine buildings. We explore this relationship from a combinatorial and computational perspective. Our results include a convex hull algorithm for the Bruhat--Tits building of SL$_d(K)$ and techniques for computing with apartments and membranes. While the original inspiration was the work of Dress and Terhalle in phylogenetics, and of Faltings, Kapranov, Keel and Tevelev in algebraic geometry, our tropical algorithms will also be applicable to problems in other fields of mathematics.
The Neighbor-Net Algorithm
Published
• View Publication
• BIB
The neighbor-joining algorithm is a popular phylogenetics method for constructing trees from dissimilarity maps. The neighbor-net algorithm is an extension of the neighbor-joining algorithm and is used for constructing split networks. We begin by describing the output of neighbor-net in terms of the tessellation of $\bar{\MM}_{0}^n(\mathbb{R})$ by associahedra. This highlights the fact that neighbor-net outputs a tree in addition to a circular ordering and we explain when the neighbor-net tree is the neighbor-joining tree. A key observation is that the tree constructed in existing implementations of neighbor-net is not a neighbor-joining tree. Next, we show that neighbor-net is a greedy algorithm for finding circular split systems of minimal balanced length. This leads to an interpretation of neighbor-net as a greedy algorithm for the traveling salesman problem. The algorithm is optimal for Kalmanson matrices, from which it follows that neighbor-net is consistent and has optimal radius 1/2. We also provide a statistical interpretation for the balanced length for a circular split system as the length based on weighted least squares estimates of the splits. We conclude with applications of these results and demonstrate the implications of our theorems for a recently published comparison of Papuan and Austronesian languages.
Asymptotic evolution of acyclic random mappings
Published
• View Publication
• BIB
An acyclic mapping from an $n$ element set into itself is a mapping $φ$ such that if $φ^k(x) = x$ for some $k$ and $x$, then $φ(x) = x$. Equivalently, $φ^\ell = φ^{\ell+1} = ...$ for $\ell$ sufficiently large. We investigate the behavior as $n \to \infty$ of a Markov chain on the collection of such mappings. At each step of the chain, a point in the $n$ element set is chosen uniformly at random and the current mapping is modified by replacing the current image of that point by a new one chosen independently and uniformly at random, conditional on the resulting mapping being again acyclic. We can represent an acyclic mapping as a directed graph (such a graph will be a collection of rooted trees) and think of these directed graphs as metric spaces with some extra structure. Heuristic calculations indicate that the metric space valued process associated with the Markov chain should, after an appropriate time and ``space'' rescaling, converge as $n \to \infty$ to a real tree ($\R$-tree) valued Markov process that is reversible with respect to a measure induced naturally by the standard reflected Brownian bridge. The limit process, which we construct using Dirichlet form methods, is a Hunt process with respect to a suitable Gromov-Hausdorff-like metric. This process is similar to one that appears in earlier work by Evans and Winter as the limit of chains involving the subtree prune and regraft tree (SPR) rearrangements from phylogenetics.
Stochastic Models for Speciation Events in Phylogenetic trees
In a phylogenetic tree, we often don't have information about the time a speciation event (inner node) occured. Under a neutral model for speciation, I develop fast algorithms for calculating the probability that an inner node i is the k-th speciation event. For the Yule and the coalescent model, I develop an edge length estimation as well. Various properties of the Yule model are discussed throughout the thesis.
Estimating the relative order of speciation or coalescence events on a given phylogeny
Published
• View Publication
• BIB
The reconstruction of large phylogenetic trees from data that violates clocklike evolution (or as a supertree constructed from any m input trees) raises a difficult question for biologists - how can one assign relative dates to the vertices of the tree? In this paper we investigate this problem, assuming a uniform distribution on the order of the inner vertices of the tree (which includes, but is more general than, the popular Yule distribution on trees). We derive fast algorithms for computing the probability that (i) any given vertex in the tree was the j--th speciation event (for each j), and (ii) any one given vertex is earlier in the tree than a second given vertex. We show how the first algorithm can be used to calculate the expected length of any given interior edge in any given tree that has been generated under either a constant-rate speciation model, or the coalescent model.