phylogenetic
434 papers tagged with this keyword
Compactifications of phylogenetic systems and species of electrical networks
We describe new spaces and maps. Our graphical map is a visual and numerical correspondence between spaces of circular electrical networks and circular planar split systems. When restricted to the planar circular electrical case, this graphical map finds the split system uniquely associated with the Kalmanson resistance distance of the dual network, matching the induced split system familiar from phylogenetics. This correspondence is extended to compactifications of the respective spaces, taking cactus networks to the cactus split systems defined herein. The graphical map preserves both network components and cactus structure, allowing an elegant enumeration of induced phylogenetic split systems via combinatorial species. We introduce the global spaces of circular planar electrical networks and circular split systems. These new spaces are also CW complexes, but the 0-cells of each are counted by the Bell numbers as opposed to the Catalan numbers. As species, the two sorts of global cacti are seen to be compositions in complementary ways.
The $B_2$ index of galled trees
In recent years, there has been an effort to extend the classical notion of phylogenetic balance, originally defined in the context of trees, to networks. One of the most natural ways to do this is with the so-called $B_2$ index. In this paper, we study the $B_2$ index for a prominent class of phylogenetic networks: galled trees. We show that the $B_2$ index of a uniform leaf-labeled galled tree converges in distribution as the network becomes large. We characterize the corresponding limiting distribution, and show that its expected value is 2.707911858984... This is the first time that a balance index has been studied to this level of detail for a random phylogenetic network.
One specificity of this work is that we use two different and independent approaches, each with its advantages: analytic combinatorics, and local limits. The analytic combinatorics approach is more direct, as it relies on standard tools; but it involves slightly more complex calculations. Because it has not previously been used to study such questions, the local limit approach requires developing an extensive framework beforehand; however, this framework is interesting in itself and can be used to tackle other similar problems.
Sackin Indices for Labeled and Unlabeled Classes of Galled Trees
Published
• View Publication
• BIB
The Sackin index is an important measure for the balance of phylogenetic trees. We investigate two extensions of the Sackin index to the class of galled trees and two of its subclasses (simplex galled trees and normal galled trees) where we consider both labeled and unlabeled galled trees. In all cases, we show that the mean of the Sackin index for a network which is uniformly sampled from its class is asymptotic to $μn^{3/2}$ for an explicit constant $μ$. In addition, we show that the scaled Sackin index convergences weakly and with all its moments to the Airy distribution.
When is a set of phylogenetic trees displayed by a normal network?
Published
• View Publication
• BIB
A normal network is uniquely determined by the set of phylogenetic trees that it displays. Given a set $\mathcal{P}$ of rooted binary phylogenetic trees, this paper presents a polynomial-time algorithm that reconstructs the unique binary normal network whose set of displayed binary trees is $\mathcal{P}$, if such a network exists. Additionally, we show that any two rooted phylogenetic trees can be displayed by a normal network and show that this result does not extend to more than two trees. This is in contrast to tree-child networks where it has been previously shown that any collection of rooted phylogenetic trees can be displayed by a tree-child network. Lastly, we introduce a type of cherry-picking sequence that characterises when a collection $\mathcal{P}$ of rooted phylogenetic trees can be displayed by a normal network and, further, characterise the minimum number of reticulations needed over all normal networks that display $\mathcal{P}$. We then exploit these sequences to show that, for all $n\ge 3$, there exist two rooted binary phylogenetic trees on $n$ leaves that can be displayed by a tree-child network with a single reticulation, but cannot be displayed by a normal network with less than $n-2$ reticulations.
Network Representation and Modular Decomposition of Combinatorial Structures: A Galled-Tree Perspective
Published
• View Publication
• BIB
In phylogenetics, reconstructing rooted trees from distances between taxa is a common task. Böcker and Dress generalized this concept by introducing symbolic dated maps $δ:X \times X \to Υ$, where distances are replaced by symbols, and showed that there is a one-to-one correspondence between symbolic ultrametrics and labeled rooted phylogenetic trees. Many combinatorial structures fall under the umbrella of symbolic dated maps, such as 2-dissimilarities, symmetric labeled 2-structures, or edge-colored complete graphs, and are here referred to as strudigrams. Strudigrams have a unique decomposition into non-overlapping modules, which can be represented by a modular decomposition tree (MDT). In the absence of prime modules, strudigrams are equivalent to symbolic ultrametrics, and the MDT fully captures the relationships $δ(x,y)$ between pairs of vertices $x,y$ in $X$ through the label of their least common ancestor in the MDT. However, in the presence of prime vertices, this information is generally hidden. To provide this missing structural information, we aim to locally replace the prime vertices in the MDT to obtain networks that capture full information about the strudigrams. While starting with the general framework of prime-vertex replacement networks, we then focus on a specific type of such networks obtained by replacing prime vertices with so-called galls, resulting in labeled galled-trees. We introduce the concept of galled-tree explainable (GATEX) strudigrams, provide their characterization, and demonstrate that recognizing these structures and reconstructing the labeled networks that explain them can be achieved in polynomial time.
Bounding the softwired parsimony score of a phylogenetic network
In comparison to phylogenetic trees, phylogenetic networks are more suitable to represent complex evolutionary histories of species whose past includes reticulation such as hybridisation or lateral gene transfer. However, the reconstruction of phylogenetic networks remains challenging and computationally expensive due to their intricate structural properties. For example, the small parsimony problem that is solvable in polynomial time for phylogenetic trees, becomes NP-hard on phylogenetic networks under softwired and parental parsimony, even for a single binary character and structurally constrained networks. To calculate the parsimony score of a phylogenetic network $N$, these two parsimony notions consider different exponential-size sets of phylogenetic trees that can be extracted from $N$ and infer the minimum parsimony score over all trees in the set. In this paper, we ask: What is the maximum difference between the parsimony score of any phylogenetic tree that is contained in the set of considered trees and a phylogenetic tree whose parsimony score equates to the parsimony score of $N$? Given a gap-free sequence alignment of multi-state characters and a rooted binary level-$k$ phylogenetic network, we use the novel concept of an informative blob to show that this difference is bounded by $k+1$ times the softwired parsimony score of $N$. In particular, the difference is independent of the alignment length and the number of character states. We show that an analogous bound can be obtained for the softwired parsimony score of semi-directed networks, while under parental parsimony on the other hand, such a bound does not hold.
Phylogenetic degrees for Jukes-Cantor model
Jukes-Cantor model is one of the most meaningful statistical models from a biological perspective. We are interested in computing the algebraic degrees for phylogenetic varieties, which we call phylogenetic degrees, associated to the Jukes-Cantor model and any tree. As these varieties are toric, their geometry is hidden in the associated polytopes. For this reason, we provide two different combinatorial approaches to compute the volume for these polytopes.
Sparsification of Phylogenetic Covariance Matrices of $k$-Regular Trees
Published
• View Publication
• BIB
Consider a tree $T=(V,E)$ with root $\circ$ and edge length function $\ell:E\to\mathbb{R}_+$. The phylogenetic covariance matrix of $T$ is the matrix $C$ with rows and columns indexed by $L$, the leaf set of $T$, with entries $C(i,j):=\sum_{e\in[i\wedge j,o]}\ell(e)$, for each $i,j\in L$. Recent work [15] has shown that the phylogenetic covariance matrix of a large, random binary tree $T$ is significantly sparsified with overwhelmingly high probability under a change-of-basis with respect to the so-called Haar-like wavelets of $T$. This finding notably enables manipulating the spectrum of covariance matrices of large binary trees without the necessity to store them in computer memory but instead performing two post-order traversals of the tree. Building on the methods of [15], this manuscript further advances their sparsification result to encompass the broader class of $k$-regular trees, for any given $k\ge2$. This extension is achieved by refining existing asymptotic formulas for the mean and variance of the internal path length of random $k$-regular trees, utilizing hypergeometric function properties and identities.
A dissimilarity measure for semidirected networks
Published
• View Publication
• BIB
Semidirected networks have received interest in evolutionary biology as the appropriate generalization of unrooted trees to networks, in which some but not all edges are directed. Yet these networks lack proper theoretical study. We define here a general class of semidirected phylogenetic networks, with a stable set of leaves, tree nodes and hybrid nodes. We prove that for these networks, if we locally choose the direction of one edge, then globally the set of directed paths starting by this edge is stable across all choices to root the network. We define an edge-based representation of semidirected phylogenetic networks and use it to define a dissimilarity between networks, which can be efficiently computed in near-quadratic time. Our dissimilarity extends the widely-used Robinson-Foulds distance on both rooted trees and unrooted trees. After generalizing the notion of tree-child networks to semidirected networks, we prove that our edge-based dissimilarity is in fact a distance on the space of tree-child semidirected phylogenetic networks.
A Vector Representation for Phylogenetic Trees
Published
• View Publication
• BIB
Good representations for phylogenetic trees and networks are important for optimizing storage efficiency and implementation of scalable methods for the inference and analysis of evolutionary trees for genes, genomes and species. We introduce a new representation for rooted phylogenetic trees that encodes a binary tree on n taxa as a vector of length 2n in which each taxon appears exactly twice. Using this new tree representation, we introduce a novel tree rearrangement operator, called a HOP, that results in a tree space of diameter n and a quadratic neighbourhood size. We also introduce a novel metric, the HOP distance, which is the minimum number of HOPs to transform a tree into another tree. The HOP distance can be computed in near-linear time, a rare instance of a tree rearrangement distance that is tractable. Our experiments show that the HOP distance is better correlated to the Subtree-Prune-and-Regraft distance than the widely used Robinson-Foulds distance. We also describe how the novel tree representation we introduce can be further generalized to tree-child networks.
Critical beta-splitting, via contraction
Published in Electron. Commun. Probab. 30, Paper No. 10, 14 p. (2025)
• View Publication
• BIB
The critical beta-splitting tree, introduced by Aldous, is a Markov branching phylogenetic tree. Aldous and Pittel recently proved, amongst other results, a central limit theorem for the height of a random leaf. We give an alternative proof, via contraction methods for random recursive structures. These methods were developed by Neininger and Rüschendorf, motivated by Pittel's article "Normal convergence problem? Two moments and a recurrence may be the clues." Aldous and Pittel estimated the leading order terms in the first two moments. More recently, Aldous and Janson obtained an asymptotic expansion for the average height. We show that a central limit theorem follows, and bound the distance to normality. Our results also apply to the continuous version of the model, in which branching times are exponential.
Getting to the Root of the Problem: Sums of Squares for Limits of Trees
Published
• View Publication
• BIB
The inducibility of a graph represents its maximum density as an induced subgraph over all possible sequences of graphs of size growing to infinity. This invariant of graphs has been extensively studied since its introduction in $1975$ by Pippenger and Golumbic. In $2017$, Czabarka, Székely and Wagner extended this notion to leaf-labeled rooted binary trees, which are objects widely studied in the field of phylogenetics. They obtain the first results and bounds for the densities and inducibilities of such trees. Following up on their work, we apply Razborov's flag algebra theory to this setting, introducing the flag algebra of rooted leaf-labeled binary trees. This framework allows us to use polynomial optimization methods, based on semidefinite programming, to efficiently obtain new upper bounds for the inducibility of trees and to improve existing ones. Additionally, we obtain the first outer approximations of profiles of trees, which represent all possible simultaneous densities of a pair of trees in a sequence of trees of growing sizes. Finally, we are able to prove the non-convexity of some of these profiles.
Phylogenetic diversity indices from an affine and projective viewpoint
Published
• View Publication
• BIB
Phylogenetic diversity indices are commonly used to rank the elements in a collection of species or populations for conservation purposes. The derivation of these indices is typically based on some quantitative description of the evolutionary history of the species in question, which is often given in terms of a phylogenetic tree. Both rooted and unrooted phylogenetic trees can be employed, and there are close connections between the indices that are derived in these two different ways. In this paper, we introduce more general phylogenetic diversity indices that can be derived from collections of subsets (clusters) and collections of bipartitions (splits) of the given set of species. Such indices could be useful, for example, in case there is some uncertainty in the topology of the tree being used to derive a phylogenetic diversity index. As well as characterizing some of the indices that we introduce in terms of their special properties, we provide a link between cluster-based and split-based phylogenetic diversity indices that uses a discrete analogue of the classical link between affine and projective geometry. This provides a unified framework for many of the various phylogenetic diversity indices used in the literature based on rooted and unrooted phylogenetic trees, generalizations and new proofs for previous results concerning tree-based indices, and a way to define some new phylogenetic diversity indices that naturally arise as affine or projective variants of each other.
Exact and Heuristic Computation of the Scanwidth of Directed Acyclic Graphs
Published
• View Publication
• BIB
To measure the tree-likeness of a directed acyclic graph (DAG), a new width parameter that considers the directions of the arcs was recently introduced: scanwidth. We present the first algorithm that efficiently computes the exact scanwidth of general DAGs. For DAGs with one root and scanwidth $k$ it runs in $O(k \cdot n^k \cdot m)$ time. The algorithm also functions as an FPT algorithm with complexity $O(2^{4 \ell - 1} \cdot \ell \cdot n + n^2)$ for phylogenetic networks of level-$\ell$, a type of DAG used to depict evolutionary relationships among species. Our algorithm performs well in practice, being able to compute the scanwidth of synthetic networks up to 30 reticulations and 100 leaves within 500 seconds. Furthermore, we propose a heuristic that obtains an average practical approximation ratio of 1.5 on these networks. While we prove that the scanwidth is bounded from below by the treewidth of the underlying undirected graph, experiments suggest that for networks the parameters are close in practice.
On the correctness of Maximum Parsimony for data with few substitutions in the NNI neighborhood of phylogenetic trees
Published
• View Publication
• BIB
Estimating phylogenetic trees, which depict the relationships between different species, from aligned sequence data (such as DNA, RNA, or proteins) is one of the main aims of evolutionary biology. However, tree reconstruction criteria like maximum parsimony do not necessarily lead to unique trees and in some cases even fail to recognize the \enquote{correct} tree (i.e., the tree on which the data was generated). On the other hand, a recent study has shown that for an alignment containing precisely those binary characters (sites) which require up to two substitutions on a given tree, this tree will be the unique maximum parsimony tree.
It is the aim of the present paper to generalize this recent result in the following sense: We show that for a tree $T$ with $n$ leaves, as long as $k<\frac{n}{8}+\frac{11}{9}-\frac{1}{18}\sqrt{9\cdot \left(\frac{n}{4}\right)^2+16}$ (or, equivalently, $n>9 k-11+\sqrt{9k^2-22 k+17} $, which in particular holds for all $n\geq 12k$), the maximum parsimony tree for the alignment containing all binary characters which require (up to or precisely) $k$ substitutions on $T$ will be unique in the NNI neighborhood of $T$ and it will coincide with $T$, too. In other words, within the NNI neighborhood of $T$, $T$ is the unique most parsimonious tree for the said alignment. This partially answers a recently published conjecture affirmatively. Additionally, we show that for $n\geq 8$ and for $k$ being in the order of $\frac{n}{2}$, there is always a pair of phylogenetic trees $T$ and $T'$ which are NNI neighbors, but for which the alignment of characters requiring precisely $k$ substitutions each on $T$ in total requires fewer substitutions on $T'$.
Identifying circular orders for blobs in phylogenetic networks
Published
• View Publication
• BIB
Interest in the inference of evolutionary networks relating species or populations has grown with the increasing recognition of the importance of hybridization, gene flow and admixture, and the availability of large-scale genomic data. However, what network features may be validly inferred from various data types under different models remains poorly understood. Previous work has largely focused on level-1 networks, in which reticulation events are well separated, and on a general network's tree of blobs, the tree obtained by contracting every blob to a node. An open question is the identifiability of the topology of a blob of unknown level. We consider the identifiability of the circular order in which subnetworks attach to a blob, first proving that this order is well-defined for outer-labeled planar blobs. For this class of blobs, we show that the circular order information from 4-taxon subnetworks identifies the full circular order of the blob. Similarly, the circular order from 3-taxon rooted subnetworks identifies the full circular order of a rooted blob. We then show that subnetwork circular information is identifiable from certain data types and evolutionary models. This provides a general positive result for high-level networks, on the identifiability of the ordering in which taxon blocks attach to blobs in outer-labeled planar networks. Finally, we give examples of blobs with different internal structures which cannot be distinguished under many models and data types.
Equidistant Circular Split Networks
Published
• View Publication
• BIB
Phylogenetic networks are generalizations of trees that allow for the modeling of non-tree like evolutionary processes. Split networks give a useful way to construct networks with intuitive distance structures induced from the associated split graph. We explore the polyhedral geometry of distance matrices built from circular split systems which have the added property of being equidistant. We give a characterization of the facet defining inequalities and the extreme rays of the cone of distances that arises from an equidistant network associated to any circular split network. We also explain a connection to the Chan-Robbins-Yuen polytope from geometric combinatorics.
0-1 laws for pattern occurrences in phylogenetic trees and networks
Published in Bull. Math. Biol. 86, 94 (2024)
• View Publication
• BIB
In a recent paper, the question of determining the fraction of binary trees that contain a fixed pattern known as the snowflake was posed. We show that this fraction goes to 1, providing two very different proofs: a purely combinatorial one that is quantitative and specific to this problem; and a proof using branching process techniques that is less explicit, but also much more general, as it applies to any fixed patterns and can be extended to other trees and networks. In particular, it follows immediately from our second proof that the fraction of $d$-ary trees (resp. level-$k$ networks) that contain a fixed $d$-ary tree (resp. level-$k$ network) tends to $1$ as the number of leaves grows.
Phylogenetic Trees and the Moduli Space of n Points on the Projective Line
This is an expository paper. The geometry of phylogenetic trees is used to present in an accessible and pleasant fashion the results of Deligne, Mumford, and Knudsen about the moduli space of n distinct points on the projective line and its compactification, the moduli space of n-pointed stable curves of genus zero.
Counting Phylogenetic Networks with Few Reticulation Vertices: Galled and Reticulation-Visible Networks
Published
• View Publication
• BIB
We give exact and asymptotic counting results for the number of galled networks and reticulation-visible networks with few reticulation vertices. Our results are obtained with the component graph method, which was introduced by L. Zhang and his coauthors, and generating function techniques. For galled networks, we in addition use analytic combinatorics. Moreover, in an appendix, we consider maximally reticulated reticulation-visible networks and derive their number, too.