arXiv++ Combinatorics

Browse math.CO papers from arXiv

phylogenetic

434 papers tagged with this keyword
2018-05-31 v8
Tropical Geometry of Phylogenetic Tree Space: A Statistical Perspective
Phylogenetic trees are the fundamental mathematical representation of evolutionary processes in biology. They are also objects of interest in pure mathematics, such as algebraic geometry and combinatorics, due to their discrete geometry. Although they are important data structures, they face the significant challenge that sets of trees form a non-Euclidean phylogenetic tree space, which means that standard computational and statistical methods cannot be directly applied. In this work, we explore the statistical feasibility of a pure mathematical representation of the set of all phylogenetic trees based on tropical geometry for both descriptive and inferential statistics, and unsupervised and supervised machine learning. Our exploration is both theoretical and practical. We show that the tropical geometric phylogenetic tree space endowed with a generalized Hilbert projective metric exhibits analytic, geometric, and topological properties that are desirable for theoretical studies in probability and statistics and allow for well-defined questions to be posed. We illustrate the statistical feasibility of the tropical geometric perspective for phylogenetic trees with an example of both a descriptive and inferential statistical task. Moreover, this approach exhibits increased computational efficiency and statistical performance over the current state-of-the-art, which we illustrate with a real data example on seasonal influenza. Our results demonstrate the viability of the tropical geometric setting for parametric statistical and probabilistic studies of sets of phylogenetic trees.
2018-05-20 v2
On the Subnet Prune and Regraft Distance
Published in The Electronic Journal of Combinatorics 26(2) (2019), #P2.3 • View Publication • BIB
Phylogenetic networks are rooted directed acyclic graphs that represent evolutionary relationships between species whose past includes reticulation events such as hybridisation and horizontal gene transfer. To search the space of phylogenetic networks, the popular tree rearrangement operation rooted subtree prune and regraft (rSPR) was recently generalised to phylogenetic networks. This new operation - called subnet prune and regraft (SNPR) - induces a metric on the space of all phylogenetic networks as well as on several widely-used network classes. In this paper, we investigate several problems that arise in the context of computing the SNPR-distance. For a phylogenetic tree $T$ and a phylogenetic network $N$, we show how this distance can be computed by considering the set of trees that are embedded in $N$ and then use this result to characterise the SNPR-distance between $T$ and $N$ in terms of agreement forests. Furthermore, we analyse properties of shortest SNPR-sequences between two phylogenetic networks $N$ and $N'$, and answer the question whether or not any of the classes of tree-child, reticulation-visible, or tree-based networks isometrically embeds into the class of all phylogenetic networks under SNPR.
2018-05-10 v2
The Cavender-Farris-Neyman Model with a Molecular Clock
Published • View Publication • BIB
We give a combinatorial description of the toric ideal of invariants of the Cavender-Farris-Neyman model with a molecular clock (CFN-MC) on a rooted binary phylogenetic tree and prove results about the polytope associated to this toric ideal. Key results about the polyhedral structure include that the number of vertices of this polytope is a Fibonacci number, the facets of the polytope can be described using the combinatorial "cluster" structure of the underlying rooted tree, and the volume is equal to an Euler zig-zag number. The toric ideal of invariants of the CFN-MC model has a quadratic Groebner basis with squarefree initial terms. Finally, we show that the Ehrhart polynomial of these polytopes, and therefore the Hilbert series of the ideals, depends only on the number of leaves of the underlying binary tree, and not on the topology of the tree itself. These results are analogous to classic results for the Cavender-Farris-Neyman model without a molecular clock. However, new techniques are required because the molecular clock assumption destroys the toric fiber product structure that governs group-based models without the molecular clock.
2018-05-03
Sound Colless-like balance indices for multifurcating trees
Published • View Publication • BIB
The Colless index is one of the most popular and natural balance indices for bifurcating phylogenetic trees, but it makes no sense for multifurcating trees. In this paper we propose a family of Colless-like balance indices $\mathfrak{C}_{D,f}$, which depend on a dissimilarity $D$ and a function $f:\mathbb{N}\to \mathbb{R}_{\geq 0}$, that generalize the Colless index to multifurcating phylogenetic trees. We provide two functions $f$ such that the most balanced phylogenetic trees according to the corresponding indices $\mathfrak{C}_{D,f}$ are exactly the fully symmetric ones. Next, for each one of these two functions $f$ and for three popular dissimilarities $D$ (the variance, the standard deviation, and the mean deviation from the median), we determine the range of values of $\mathfrak{C}_{D,f}$ on the sets of phylogenetic trees with a given number $n$ of leaves. We end the paper by assessing the performance of one of these indices on TreeBASE and using it to show that the trees in this database do not seem to follow either the uniform model for multifurcating trees or the $α$-$γ$-model, for any values of $α$ and $γ$.
2018-04-05
The polytopal structure of the tight-span of a totally split-decomposable metric
Published • View Publication • BIB
The tight-span of a finite metric space is a polytopal complex that has appeared in several areas of mathematics. In this paper we determine the polytopal structure of the tight-span of a totally split decomposable (finite) metric. Totally split-decomposable metrics are a generalization of tree-metrics and have importance within phylogenetics. In previous work, we showed that the cells of the tight-span of such a metric are zonotopes that are polytope isomorphic to either hypercubes or rhombic dodecahedra. Here, we extend these results and show that the tight-spanof a totally split-decomposable metric can be broken up into a canonical collection of polytopal complexes whose polytopal structures can be directly determined from the metric. This allows us to also completely determine the polytopal structure of the tight-span of a totally split-decomposable metric in a very direct way.We anticipate that our improved understanding of this structure may ultimately lead to improved techniques for phylogenetic inference.
2018-04-05 v3
Phylogenetic networks that are their own fold-ups
Published • View Publication • BIB
Phylogenetic networks are becoming of increasing interest to evolutionary biologists due to their ability to capture complex non-treelike evolutionary processes. From a combinatorial point of view, such networks are certain types of rooted directed acyclic graphs whose leaves are labelled by, for example, species. A number of mathematically interesting classes of phylogenetic networks are known. These include the biologically relevant class of stable phylogenetic networks whose members are defined via certain "fold-up" and "un-fold" operations that link them with concepts arising within the theory of, for example, graph fibrations. Despite this exciting link, the structural complexity of stable phylogenetic networks is still relatively poorly understood. Employing the popular tree-based, reticulation-visible, and tree-child properties which allow one to gauge this complexity in one way or another, we provide novel characterizations for when a stable phylogenetic network satisfies either one of these three properties.
Counting Phylogenetic Networks with Few Reticulation Vertices: Tree-Child and Normal Networks
In recent decades, phylogenetic networks have become a standard tool in modeling evolutionary processes. Nevertheless, basic combinatorial questions about them are still largely open. For instance, even the asymptotic counting problem for the class of phylogenetic networks and subclasses is unsolved. In this paper, we propose a method based on generating functions to count networks with few reticulation vertices for two subclasses which are important in applications: tree-child networks and normal networks. In particular, our method can be used to completely solve the asymptotic counting problem for these network classes when the number of reticulation vertices remains fixed and the network size tends to infinity.
Best Match Graphs
Published • View Publication • BIB
THIS IS A CORRECTED VERSION INCLUDING AN APPENDED CORRIGENDUM. Best match graphs arise naturally as the first processing intermediate in algorithms for orthology detection. Let $T$ be a phylogenetic (gene) tree $T$ and $σ$ an assignment of leaves of $T$ to species. The best match graph $(G,σ)$ is a digraph that contains an arc from $x$ to $y$ if the genes $x$ and $y$ reside in different species and $y$ is one of possibly many (evolutionary) closest relatives of $x$ compared to all other genes contained in the species $σ(y)$. Here, we characterize best match graphs and show that it can be decided in cubic time and quadratic space whether $(G,σ)$ derived from a tree in this manner. If the answer is affirmative, there is a unique least resolved tree that explains $(G,σ)$, which can also be constructed in cubic time.
2018-03-16 v3
Shellability of face posets of electrical networks and the CW poset property
Published in Advances in Applied Math, 127 (2021), 37 pages • View Publication • BIB
We prove a conjecture of Thomas Lam that the face posets of stratified spaces of planar resistor networks are shellable. These posets are called uncrossing partial orders. This shellability result combines with Lam's previous result that these same posets are Eulerian to imply that they are CW posets, namely that they are face posets of regular CW complexes. Certain subsets of uncrossing partial orders are shown to be isomorphic to type A Bruhat order intervals; our shelling is shown to coincide on these intervals with a Bruhat order shelling which was constructed by Matthew Dyer using a reflection order. Our shelling for uncrossing posets also yields an explicit shelling for each interval in the face posets of the edge product spaces of phylogenetic trees, namely in the Tuffley posets, by virtue of each interval in a Tuffley poset being isomorphic to an interval in an uncrossing poset. This yields a more explicit proof of the result of Gill, Linusson, Moulton and Steel that the CW decomposition of Moulton and Steel for the edge product space of phylogenetic trees is a regular CW decomposition.
The complexity of comparing multiply-labelled trees by extending phylogenetic-tree metrics
Published • View Publication • BIB
A multilabeled tree (or MUL-tree) is a rooted tree in which every leaf is labelled by an element from some set, but in which more than one leaf may be labelled by the same element of that set. In phylogenetics, such trees are used in biogeographical studies, to study the evolution of gene families, and also within approaches to construct phylogenetic networks. A multilabelled tree in which no leaf-labels are repeated is called a phylogenetic tree, and one in which every label is the same is also known as a tree-shape. In this paper, we consider the complexity of computing metrics on MUL-trees that are obtained by extending metrics on phylogenetic trees. In particular, by restricting our attention to tree shapes, we show that computing the metric extension on MUL-trees is NP complete for two well-known metrics on phylogenetic trees, namely, the path-difference and Robinson Foulds distances. We also show that the extension of the Robinson Foulds distance is fixed parameter tractable with respect to the distance parameter. The path distance complexity result allows us to also answer an open problem concerning the complexity of solving the quadratic assignment problem for two matrices that are a Robinson similarity and a Robinson dissimilarity, which we show to be NP-complete. We conclude by considering the maximum agreement subtree (MAST) distance on phylogenetic trees to MUL-trees. Although its extension to MUL-trees can be computed in polynomial time, we show that computing its natural generalization to more than two MUL-trees is NP-complete, although fixed-parameter tractable in the maximum degree when the number of given trees is bounded.
2018-03-08
Not all phylogenetic networks are leaf-reconstructible
Published • View Publication • BIB
Unrooted phylogenetic networks are graphs used to represent evolutionary relationships. Accurately reconstructing such networks is of great relevance for evolutionary biology. It has recently been conjectured that all phylogenetic networks with at least five leaves can be uniquely reconstructed from their subnetworks obtained by deleting a single leaf and suppressing degree-2 vertices. Here, we show that this conjecture is false, by presenting a counter example for each possible number of leaves that is at least~4. Moreover, we show that the conjecture is still false when restricted to binary networks.
2018-02-10 v2
Generalized Fitch Graphs: Edge-labeled Graphs that are explained by Edge-labeled Trees
Published • View Publication • BIB
Fitch graphs $G=(X,E)$ are di-graphs that are explained by $\{\otimes,1\}$-edge-labeled rooted trees with leaf set $X$: there is an arc $xy\in E$ if and only if the unique path in $T$ that connects the least common ancestor $\textrm{lca}(x,y)$ of $x$ and $y$ with $y$ contains at least one edge with label $1$. In practice, Fitch graphs represent xenology relations, i.e., pairs of genes $x$ and $y$ for which a horizontal gene transfer happened along the path from $\textrm{lca}(x,y)$ to $y$. In this contribution, we generalize the concept of xenology and Fitch graphs and consider complete di-graphs $K_{|X|}$ with vertex set $X$ and a map $ε$ that assigns to each arc $xy$ a unique label $ε(x,y)\in M\cup \{\otimes\}$, where $M$ denotes an arbitrary set of symbols. A di-graph $(K_{|X|},ε)$ is a generalized Fitch graph if there is an $M\cup \{\otimes\}$-edge-labeled tree $(T,λ)$ that can explain $(K_{|X|},ε)$. We provide a simple characterization of generalized Fitch graphs $(K_{|X|},ε)$ and give an $O(|X|^2)$-time algorithm for their recognition as well as for the reconstruction of the unique least resolved phylogenetic tree that explains $(K_{|X|},ε)$.
2018-02-07 v5
Combinatorial views on persistent characters in phylogenetics
Published • View Publication • BIB
The so-called binary perfect phylogeny with persistent characters has recently been thoroughly studied in computational biology as it is less restrictive than the well known binary perfect phylogeny. Here, we focus on the notion of (binary) persistent characters, i.e. characters that can be realized on a phylogenetic tree by at most one $0 \rightarrow 1$ transition followed by at most one $1 \rightarrow 0$ transition in the tree, and analyze these characters under different aspects. First, we illustrate the connection between persistent characters and Maximum Parsimony, where we characterize persistent characters in terms of the first phase of the famous Fitch algorithm. Afterwards we focus on the number of persistent characters for a given phylogenetic tree. We show that this number solely depends on the balance of the tree. To be precise, we develop a formula for counting the number of persistent characters for a given phylogenetic tree based on an index of tree balance, namely the Sackin index. Lastly, we consider the question of how many (carefully chosen) binary characters together with their persistence status are needed to uniquely determine a phylogenetic tree and provide an upper bound for the number of characters needed.
2018-01-31 v6
Extremal values of the Sackin tree balance index
Published • View Publication • BIB
Tree balance plays an important role in different research areas like theoretical computer science and mathematical phylogenetics. For example, it has long been known that under the Yule model, a pure birth process, imbalanced trees are more likely than balanced ones. Also, concerning ordered search trees, more balanced ones allow for more efficient data structuring than imbalanced ones. Therefore, different methods to measure the balance of trees were introduced. The Sackin index is one of the most frequently used measures for this purpose. In many contexts, statements about the minimal and maximal values of this index have been discussed, but formal proofs have only been provided for some of them, and only in the context of ordered binary (search) trees, not for general rooted trees. Moreover, while the number of trees with maximal Sackin index as well as the number of trees with minimal Sackin index when the number of leaves is a power of 2 are relatively easy to understand, the number of trees with minimal Sackin index for all other numbers of leaves has been completely unknown. In this manuscript, we extend the findings on trees with minimal and maximal Sackin indices from the literature on ordered trees and subsequently use our results to provide formulas to explicitly calculate the numbers of such trees. We also extend previous studies by analyzing the case when the underlying trees need not be binary. Finally, we use our results to contribute both to the phylogenetic as well as the computer scientific literature by using the new findings on Sackin minimal and maximal trees in order to derive formulas to calculate the number of both minimal and maximal phylogenetic trees as well as minimal and maximal ordered trees both in the binary and non-binary settings. All our results have been implemented in the Mathematica package SackinMinimizer, which has been made publicly available.
2018-01-02
Phylogenetic trees and homomorphisms
In Chapter 1 we fully characterise pairs of finite graphs which form a gap in the full homomorphism order. This leads to a simple proof of the existence of generalised duality pairs. We also discuss how such results can be carried to relational structures with unary and binary relations. In Chapter 2 we show a very simple and versatile argument based on divisibility which immediately yields the universality of the homomorphism order of directed graphs and discuss three applications. In chapter 3, we show that every interval in the homomorphism order of finite undirected graphs is either universal or a gap. Together with density and universality this "fractal" property contributes to the spectacular properties of the homomorphism order. In Chapter 4 we analyze the phylogenetic information content from a combinatorial point of view by considering the binary relation on the set of taxa defined by the existence of a single event separating two taxa. We show that the graph-representation of this relation must be a tree. Moreover, we characterize completely the relationship between the tree of such relations and the underlying phylogenetic tree.
2017-12-12 v2
Attaching leaves and picking cherries to characterise the hybridisation number for a set of phylogenies
Published in Advances in Applied Mathematics, 105:102-129, 2019 • View Publication • BIB
Throughout the last decade, we have seen much progress towards characterising and computing the minimum hybridisation number for a set P of rooted phylogenetic trees. Roughly speaking, this minimum quantifies the number of hybridisation events needed to explain a set of phylogenetic trees by simultaneously embedding them into a phylogenetic network. From a mathematical viewpoint, the notion of agreement forests is the underpinning concept for almost all results that are related to calculating the minimum hybridisation number for when |P|=2. However, despite various attempts, characterising this number in terms of agreement forests for |P|>2 remains elusive. In this paper, we characterise the minimum hybridisation number for when P is of arbitrary size and consists of not necessarily binary trees. Building on our previous work on cherry-picking sequences, we first establish a new characterisation to compute the minimum hybridisation number in the space of tree-child networks. Subsequently, we show how this characterisation extends to the space of all rooted phylogenetic networks. Moreover, we establish a particular hardness result that gives new insight into some of the limitations of agreement forests.
A Short Note on Undirected Fitch Graphs
The symmetric version of Fitch's xenology relation coincides with class of complete multipartite graph and thus cannot convey any non-trivial phylogenetic information.
Recovering tree-child networks from shortest inter-taxa distance information
Published • View Publication • BIB
Phylogenetic networks are a type of leaf-labelled, acyclic, directed graph used by biologists to represent the evolutionary history of species whose past includes reticulation events. A phylogenetic network is tree-child if each non-leaf vertex is the parent of a tree vertex or a leaf. Up to a certain equivalence, it has been recently shown that, under two different types of weightings, edge-weighted tree-child networks are determined by their collection of distances between each pair of taxa. However, the size of these collections can be exponential in the size of the taxa set. In this paper, we show that, if we ignore redundant edges, the same results are obtained with only a quadratic number of inter-taxa distances by using the shortest distance between each pair of taxa. The proofs are constructive and give cubic-time algorithms in the size of the taxa sets for building such weighted networks.
Quarnet inference rules for level-1 networks
Published • View Publication • BIB
An important problem in phylogenetics is the construction of phylogenetic trees. One way to approach this problem, known as the supertree method, involves inferring a phylogenetic tree with leaves consisting of a set $X$ of species from a collection of trees, each having leaf-set some subset of $X$. In the 1980's characterizations, certain inference rules were given for when a collection of 4-leaved trees, one for each 4-element subset of $X$, can all be simultaneously displayed by a single supertree with leaf-set $X$. Recently, it has become of interest to extend such results to phylogenetic networks. These are a generalization of phylogenetic trees which can be used to represent reticulate evolution (where species can come together to form a new species). It has been shown that a certain type of phylogenetic network, called a level-1 network, can essentially be constructed from 4-leaved trees. However, the problem of providing appropriate inference rules for such networks remains unresolved. Here we show that by considering 4-leaved networks, called quarnets, as opposed to 4-leaved trees, it is possible to provide such rules. In particular, we show that these rules can be used to characterize when a collection of quarnets, one for each 4-element subset of $X$, can all be simultaneously displayed by a level-1 network with leaf-set $X$. The rules are an intriguing mixture of tree inference rules, and an inference rule for building up a cyclic ordering of $X$ from orderings on subsets of $X$ of size 4. This opens up several new directions of research for inferring phylogenetic networks from smaller ones, which could yield new algorithms for solving the supernetwork problem in phylogenetics.
2017-11-14 v2
Tree-Based Unrooted Nonbinary Phylogenetic Networks
Published • View Publication • BIB
Phylogenetic networks are a generalisation of phylogenetic trees that allow for more complex evolutionary histories that include hybridisation-like processes. It is of considerable interest whether a network can be considered `tree-like' or not, which lead to the introduction of \textit{tree-based} networks in the rooted, binary context. Tree-based networks are those networks which can be constructed by adding additional edges into a given phylogenetic tree, called the \textit{base tree}. Previous extensions have considered extending to the binary, unrooted case and the nonbinary, rooted case. We extend tree-based networks to the context of unrooted, nonbinary networks in three ways, depending on the types of additional edges that are permitted. A phylogenetic network in which every embedded tree is a base tree is termed a \textit{fully tree-based} network. We also extend this concept to unrooted, nonbinary phylogenetic networks and classify the resulting networks. We also derive some results on the colourability of tree-based networks, which can be useful to determine whether a network is tree-based.