phylogenetic
434 papers tagged with this keyword
Optimal distance query reconstruction for graphs without long induced cycles
Given access to the vertex set $V$ of a connected graph $G=(V,E)$ and an oracle that given two vertices $u,v\in V$, returns the shortest path distance between $u$ and $v$, how many queries are needed to reconstruct $E$?
Firstly, we show that randomised algorithms need to use at least $\frac1{200} Δn\log_Δn$ queries in expectation in order to reconstruct $n$-vertex trees of maximum degree $Δ$. The best previous lower bound (for graphs of bounded maximum degree) was an information-theoretic lower bound of $Ω(n\log n/\log \log n)$. Our randomised lower bound is also the first to break through the information-theoretic barrier for related query models including distance queries for phylogenetic trees, membership queries for learning partitions and path queries in directed trees.
Secondly, we provide a simple deterministic algorithm to reconstruct trees using $Δn\log_Δn+(Δ+2)n$ distance queries. This proves that our lower bound is optimal up to a multiplicative constant. We extend our algorithm to reconstruct graphs without induced cycles of length at least $k$ using $O_{Δ,k}(n\log n)$ queries. Our lower bound is therefore tight for a wide range of tree-like graphs, such as chordal graphs, permutation graphs and AT-free graphs. The previously best randomised algorithm for chordal graphs used $O_Δ(n\log^2 n)$ queries in expectation, so we improve by a $(\log n)$-factor for this graph class.
Rooted Almost-binary Phylogenetic Networks for which the Maximum Covering Subtree Problem is Solvable in Linear Time
Phylogenetic networks are a flexible model of evolution that can represent reticulate evolution and handle complex data. Tree-based networks, which are phylogenetic networks that have a spanning tree with the same root and leaf-set as the network itself, have been well studied. However, not all networks are tree-based. Francis-Semple-Steel (2018) thus introduced several indices to measure the deviation of rooted binary phylogenetic networks $N$ from being tree-based, such as the minimum number $δ^\ast(N)$ of additional leaves needed to make $N$ tree-based, and the minimum difference $η^\ast(N)$ between the number of vertices of $N$ and the number of vertices of a subtree of $N$ that shares the root and leaf set with $N$. Hayamizu (2021) has established a canonical decomposition of almost-binary phylogenetic networks of $N$, called the maximal zig-zag trail decomposition, which has many implications including a linear time algorithm for computing $δ^\ast(N)$. The Maximum Covering Subtree Problem (MCSP) is the problem of computing $η^\ast(N)$, and Davidov et al. (2022) showed that this can be solved in polynomial time (in cubic time when $N$ is binary) by an algorithm for the minimum cost flow problem. In this paper, under the assumption that $N$ is almost-binary (i.e. each internal vertex has in-degree and out-degree at most two), we show that $δ^\ast(N)\leq η^\ast (N)$ holds, which is tight, and give a characterisation of such phylogenetic networks $N$ that satisfy $δ^\ast(N)=η^\ast(N)$. Our approach uses the canonical decomposition of $N$ and focuses on how the maximal W-fences (i.e. the forbidden subgraphs of tree-based networks) are connected to maximal M-fences in the network $N$. Our results introduce a new class of phylogenetic networks for which MCSP can be solved in linear time, which can be seen as a generalisation of tree-based networks.
Orienting undirected phylogenetic networks to tree-child network
Phylogenetic networks are used to represent the evolutionary history of species. They are versatile when compared to traditional phylogenetic trees, as they capture more complex evolutionary events such as hybridization and horizontal gene transfer. Distance-based methods such as the Neighbor-Net algorithm are widely used to compute phylogenetic networks from data. However, the output is necessarily an undirected graph, posing a great challenge to deduce the direction of genetic flow in order to infer the true evolutionary history. Recently, Huber et al. investigated two different computational problems relevant to orienting undirected phylogenetic networks into directed ones. In this paper, we consider the problem of orienting an undirected binary network into a tree-child network. We give some necessary conditions for determining the tree-child orientability, such as a tight upper bound on the size of tree-child orientable graphs, as well as many interesting examples. In addition, we introduce new families of undirected phylogenetic networks, the jellyfish graphs and ladder graphs, that are orientable but not tree-child orientable. We also prove that any ladder graph can be made tree-child orientable by adding extra leaves, and describe a simple algorithm for orienting a ladder graph to a tree-child network with the minimum number of extra leaves. We pose many open problems as well.
Making a Network Orchard by Adding Leaves
Phylogenetic networks are used to represent the evolutionary history of species. Recently, the new class of orchard networks was introduced, which were later shown to be interpretable as trees with additional horizontal arcs. This makes the network class ideal for capturing evolutionary histories that involve horizontal gene transfers. Here, we study the minimum number of additional leaves needed to make a network orchard. We demonstrate that computing this proximity measure for a given network is NP-hard and describe a tight upper bound. We also give an equivalent measure based on vertex labellings to construct a mixed integer linear programming formulation. Our experimental results, which include both real-world and synthetic data, illustrate the effectiveness of our implementation.
The Theory of Gene Family Histories
Published
• View Publication
• BIB
Most genes are part of larger families of evolutionary related genes. The history of gene families typically involves duplications and losses of genes as well as horizontal transfers into other organisms. The reconstruction of detailed gene family histories, i.e., the precise dating of evolutionary events relative to phylogenetic tree of the underlying species has remained a challenging topic despite their importance as a basis for detailed investigations into adaptation and functional evolution of individual members of the gene family. The identification of orthologs, moreover, is a particularly important subproblem of the more general setting considered here. In the last few years, an extensive body of mathematical results has appeared that tightly links orthology, a formal notion of best matches among genes, and horizontal gene transfer. The purpose of this chapter is the broadly outline some of the key mathematical insights and to discuss their implication for practical applications. In particular, we focus on tree-free methods, i.e., methods to infer orthology or horizontal gene transfer as well as gene trees, species trees and reconciliations between them without using \emph{a priori} knowledge of the underlying trees or statistical models for the inference of phylogenetic trees. Instead, the initial step aims to extract binary relations among genes.
Quantifying the difference between phylogenetic diversity and diversity indices
Published
• View Publication
• BIB
Phylogenetic diversity is a popular measure for quantifying the biodiversity of a collection $Y$ of species, while phylogenetic diversity indices provide a way to apportion phylogenetic diversity to individual species. Typically, for some specific diversity index, the phylogenetic diversity of $Y$ is not equal to the sum of the diversity indices of the species in $Y.$ In this paper, we investigate the extent of this difference for two commonly-used indices: Fair Proportion and Equal Splits. In particular, we determine the maximum value of this difference under various instances including when the associated rooted phylogenetic tree is allowed to vary across all root phylogenetic trees with the same leaf set and whose edge lengths are constrained by either their total sum or their maximum value.
Phylogenetic degrees for claw trees
Published
• View Publication
• BIB
Group-based models appear in algebraic statistics as mathematical models coming from evolutionary biology, respectively the study of mutations of organisms. Both theoretically and in terms of applications, we are interested in determining the algebraic degrees of the phylogenetic varieties coming from these models. These algebraic degrees are called phylogenetic degrees. In this paper, we compute the phylogenetic degree of the variety $X_{G, n}$ with $G\in\{\mathbb{Z}_2,\mathbb{Z}_2\times\mathbb{Z}_2, \mathbb{Z}_3\}$ and any $n$-claw tree. As these varieties are toric, computing their phylogenetic degree relies on computing the volume of their associated polytopes $P_{G,n}$. We apply combinatorial methods and we give concrete formulas for them.
Identifiability of the Rooted Tree Parameter under the Cavender-Farris-Neyman Model with a Molecular Clock
Identifiability of the discrete tree parameter is a key property for phylogenetic models since it is necessary for statistically consistent estimation of the tree from sequence data. Algebraic methods have proven to be very effective at showing that tree and network parameters of phylogenetic models are identifiable, especially when the underlying models are group-based. However, since group-based models are time-reversible, only the unrooted tree topology is identifiable and the location of the root is not. In this note we show that the rooted tree parameter of the Cavender-Farris-Neyman Model with a Molecular Clock is generically identifiable by using the invariants of the model which were characterized by Coons and Sullivant.
Defining binary phylogenetic trees using parsimony: new bounds
Published
• View Publication
• BIB
Phylogenetic trees are frequently used to model evolution. Such trees are typically reconstructed from data like DNA, RNA, or protein alignments using methods based on criteria like maximum parsimony (amongst others). Maximum parsimony has been assumed to work well for data with only few state changes. Recently, some progress has been made to formally prove this assertion. For instance, it has been shown that each binary phylogenetic tree $T$ with $n \geq 20k$ leaves is uniquely defined by the set $A_k(T)$, which consists of all characters with parsimony score $k$ on $T$. In the present manuscript, we show that the statement indeed holds for all $n \geq 4k$, thus drastically lowering the lower bound for $n$ from $20k$ to $4k$. However, it has been known that for $n \leq 2k$ and $k \geq 3$, it is not generally true that $A_k(T)$ defines $T$. We improve this result by showing that the latter statement can be extended from $n \leq 2k$ to $n \leq 2k+2$. So we drastically reduce the gap of values of $n$ for which it is unknown if trees $T$ on $n$ taxa are defined by $A_k(T)$ from the previous interval of $[2k+1,20k-1]$ to the interval $[2k+3,4k-1]$. Moreover, we close this gap completely for the nearest neighbor interchange (NNI) neighborhood of $T$ in the following sense: We show that as long as $n\geq 2k+3$, no tree that is one NNI move away from $T$ (and thus very similar to $T$) shares the same $A_k$-alignment.
Snakes and Ladders: a Treewidth Story
Published
• View Publication
• BIB
Let $G$ be an undirected graph. We say that $G$ contains a ladder of length $k$ if the $2 \times (k+1)$ grid graph is an induced subgraph of $G$ that is only connected to the rest of $G$ via its four cornerpoints. We prove that if all the ladders contained in $G$ are reduced to length 4, the treewidth remains unchanged (and that this bound is tight). Our result indicates that, when computing the treewidth of a graph, long ladders can simply be reduced, and that minimal forbidden minors for bounded treewidth graphs cannot contain long ladders. Our result also settles an open problem from algorithmic phylogenetics: the common chain reduction rule, used to simplify the comparison of two evolutionary trees, is treewidth-preserving in the display graph of the two trees.
Classifying Tree Topologies along Tropical Line Segments
Published in Alg. Stat. 14 (2023) 71-90
• View Publication
• BIB
The space of phylogenetic trees arises naturally in tropical geometry as the tropical Grassmannian. Tropical geometry therefore suggests a natural notion of a tropical path between two trees, given by a tropical line segment in the tropical Grassmannian. It was previously conjectured that tree topologies along such a segment change by a combinatorial operation known as Nearest Neighbor Interchange (NNI). We provide counterexamples to this conjecture, but prove that changes in tree topologies along the tropical line segment are either NNI moves or "four clade rearrangement" moves for generic trees. In addition, we show that the number of NNI moves occurring along the tropical line segment can be as large as $n^2$, but the average number of moves when the two endpoint trees are chosen at random is $O(n (\log n)^4)$. This is in contrast with $O(n \log n)$, the average number of NNI moves needed to transform one tree into another.
Exploring spaces of semi-directed phylogenetic networks
Published
• View Publication
• BIB
Semi-directed phylogenetic networks have recently emerged as a class of phylogenetic networks sitting between rooted (directed) and unrooted (undirected) phylogenetic networks as they contain both directed as well as undirected edges. While the spaces of rooted phylogenetic networks and unrooted phylogenetic networks have been analyzed in recent years and various rearrangement moves to traverse these spaces have been introduced, the results do not immediately carry over to semi-directed phylogenetic networks. Here, we propose a simple rearrangement move for semi-directed phylogenetic networks, called cut edge transfer (CET), and show that the space of semi-directed level-$1$ networks with precisely $k$ reticulations is connected under CET. This level-$1$ space is currently the predominantly used search space for most algorithms that reconstruct semi-directed phylogenetic networks. Hence, every semi-directed level-$1$ network with a fixed number of reticulations and leaf set can be reached from any other such network by a sequence of CETs. By introducing two additional moves, CET$^+$ and CET$^-$, that allow for the addition or deletion of reticulations, we then establish connectedness for the space of all semi-directed phylogenetic networks on a fixed leaf set. As a byproduct of our results for semi-directed phylogenetic networks, we also show that the space of rooted level-$1$ networks with a fixed number of reticulations and leaf set is connected under CET, when translated into the rooted setting.
Cherry picking in forests: A new characterization for the unrooted hybrid number of two phylogenetic trees
Published in Discrete Mathematics & Theoretical Computer Science, vol. 27:2, Graph Theory (May 20, 2025) dmtcs:11633
• View Publication
• BIB
Phylogenetic networks are a special type of graph which generalize phylogenetic trees and that are used to model non-treelike evolutionary processes such as recombination and hybridization. In this paper, we consider {\em unrooted} phylogenetic networks, i.e. simple, connected graphs $\mathcal{N}=(V,E)$ with leaf set $X$, for $X$ some set of species, in which every internal vertex in $\mathcal{N}$ has degree three. One approach used to construct such phylogenetic networks is to take as input a collection $\mathcal{P}$ of phylogenetic trees and to look for a network $\mathcal{N}$ that contains each tree in $\mathcal{P}$ and that minimizes the quantity $r(\mathcal{N}) = |E|-(|V|-1)$ over all such networks. Such a network always exists, and the quantity $r(\mathcal{N})$ for an optimal network $\mathcal{N}$ is called the hybrid number of $\mathcal{P}$. In this paper, we give a new characterization for the hybrid number in case $\mathcal{P}$ consists of two trees. This characterization is given in terms of a cherry picking sequence for the two trees, although to prove that our characterization holds we need to define the sequence more generally for two forests. Cherry picking sequences have been intensively studied for collections of rooted phylogenetic trees, but our new sequences are the first variant of this concept that can be applied in the unrooted setting. Since the hybrid number of two trees is equal to the well-known tree bisection and reconnection distance between the two trees, our new characterization also provides an alternative way to understand this important tree distance.
Hypercubes and Hamilton cycles of display sets of rooted phylogenetic networks
Published in Advances in Applied Mathematics, 152:102595, 2024
• View Publication
• BIB
In the context of reconstructing phylogenetic networks from a collection of phylogenetic trees, several characterisations and subsequently algorithms have been established to reconstruct a phylogenetic network that collectively embeds all trees in the input in some minimum way. For many instances however, the resulting network also embeds additional phylogenetic trees that are not part of the input. However, little is known about these inferred trees. In this paper, we explore the relationships among all phylogenetic trees that are embedded in a given phylogenetic network. First, we investigate some combinatorial properties of the collection P of all rooted binary phylogenetic trees that are embedded in a rooted binary phylogenetic network N. To this end, we associated a particular graph G, which we call rSPR graph, with the elements in P and show that, if |P|=2^k, where k is the number of vertices with in-degree two in N, then G has a Hamiltonian cycle. Second, by exploiting rSPR graphs and properties of hypercubes, we turn to the well-studied class of rooted binary level-1 networks and give necessary and sufficient conditions for when a set of rooted binary phylogenetic trees can be embedded in a level-1 network without inferring any additional trees. Lastly, we show how these conditions translate into a polynomial-time algorithm to reconstruct such a network if it exists.
Parametric Fermat-Weber and tropical supertrees
We study a parametric version of the Fermat-Weber problem with respect to an asymmetric distance function, which occurs naturally in tropical geometry. Our results yield a method for constructing phylogenetic supertrees.
Polynomial invariants for cactuses
Published
• View Publication
• BIB
Graph invariants are a useful tool in graph theory. Not only do they encode useful information about the graphs to which they are associated, but complete invariants can be used to distinguish between non-isomorphic graphs. Polynomial invariants for graphs such as the well-known Tutte polynomial have been studied for several years, and recently there has been interest to also define such invariants for phylogenetic networks, a special type of graph that arises in the area of evolutionary biology. Recently Liu gave a complete invariant for (phylogenetic) trees. However, the polynomial invariants defined thus far for phylogenetic networks that are not trees require vertex labels and either contain a large number of variables, or they have exponentially many terms in the number of reticulations. This can make it difficult to compute these polynomials and to use them to analyse unlabelled networks. In this paper, we shall show how to circumvent some of these difficulties for rooted cactuses and cactuses. As well as being important in other areas such as operations research, rooted cactuses contain some common classes of phylogenetic networks such phylogenetic trees and level-1 networks. More specifically, we define a polynomial $F$ that is a complete invariant for the class of rooted cactuses without vertices of indegree 1 and outdegree 1 that has 5 variables, and a polynomial $Q$ that is a complete invariant for the class of rooted cactuses that has 6 variables \vince{whose degree can be bounded linearly in terms of the size of the rooted cactus}. We also explain how to extend the $Q$ polynomial to define a complete invariant for leaf-labelled rooted cactuses as well as (unrooted) cactuses.
Deep kernelization for the Tree Bisection and Reconnnect (TBR) distance in phylogenetics
Published
• View Publication
• BIB
We describe a kernel of size 9k-8 for the NP-hard problem of computing the Tree Bisection and Reconnect (TBR) distance k between two unrooted binary phylogenetic trees. We achieve this by extending the existing portfolio of reduction rules with three novel new reduction rules. Two of the rules are based on the idea of topologically transforming the trees in a distance-preserving way in order to guarantee execution of earlier reduction rules. The third rule extends the local neighbourhood approach introduced in (Kelk and Linz, Annals of Combinatorics 24(3), 2020) to more global structures, allowing new situations to be identified when deletion of a leaf definitely reduces the TBR distance by one. The bound on the kernel size is tight up to an additive term. Our results also apply to the equivalent problem of computing a Maximum Agreement Forest (MAF) between two unrooted binary phylogenetic trees. We anticipate that our results will be more widely applicable for computing agreement-forest based dissimilarity measures.
Tropical medians by transportation
Published in Mathematical Programming, Volume 205 (2024), pp. 813-839
• View Publication
• BIB
Fermat-Weber points with respect to an asymmetric tropical distance function are studied. It turns out that they correspond to the optimal solutions of a transportation problem. The results are applied to obtain a new method for computing consensus trees in phylogenetics. This method has several desirable properties; e.g., it is Pareto and co-Pareto on rooted triplets.
Clustering Systems of Phylogenetic Networks
Published
• View Publication
• BIB
Rooted acyclic graphs appear naturally when the phylogenetic relationship of a set $X$ of taxa involves not only speciations but also recombination, horizontal transfer, or hybridization, that cannot be captured by trees. A variety of classes of such networks have been discussed in the literature, including phylogenetic, level-1, tree-child, tree-based, galled tree, regular, or normal networks as models of different types of evolutionary processes. Clusters arise in models of phylogeny as the sets $\mathtt{C}(v)$ of descendant taxa of a vertex $v$. The clustering system $\mathscr{C}_N$ comprising the clusters of a network $N$ conveys key information on $N$ itself. In the special case of rooted phylogenetic trees, $T$ is uniquely determined by its clustering system $\mathscr{C}_T$. Although this is no longer true for networks in general, it is of interest to relate properties of $N$ and $\mathscr{C}_N$. Here, we systematically investigate the relationships of several well-studied classes of networks and their clustering systems. The main results are correspondences of classes of networks and clustering system of the following form: If $N$ is a network of type $\mathbb{X}$, then $\mathcal{C}_N$ satisfies $\mathbb{Y}$, and conversely if $\mathscr{C}$ is a clustering system satisfying $\mathbb{Y}$ then there is network $N$ of type $\mathbb{X}$ such that $\mathscr{C}\subseteq\mathscr{C}_N$.This, in turn, allows us to investigate the mutual dependencies between the distinct types of networks in much detail.
Planar Rooted Phylogenetic Networks
Published
• View Publication
• BIB
A rooted phylogenetic network is a directed acyclic graph with a single root, whose sinks correspond to a set of species. As such networks are useful for representing the evolution of species that have undergone reticulate evolution, there has been great interest in developing the theory behind and algorithms for constructing them. However, unlike evolutionary trees, these networks can be highly non-planar, which can make them difficult to visualise and interpret. Here we investigate properties of planar rooted phylogenetic networks and algorithms for deciding whether or not rooted networks have certain special planarity properties. In particular, we introduce three natural subclasses of planar rooted phylogenetic networks and show that they form a hierarchy. In addition, for the well-known level-k networks, we show that level-1, -2, -3 networks are always outer, terminal, and upward planar, respectively, and that level-4 networks are not necessarily planar. Finally, we show that a regular network is terminal planar if and only if it is pyramidal. Our results make use of the highly developed field of planar digraphs, and we believe that the link between phylogenetic networks and planar graphs should prove useful in future for developing new approaches to both construct and visualise phylogenetic networks.