Papers by Katharina T. Huber
28 paper(s) by this author
· All BibTeX
Transforming phylogenetic networks: Moving beyond tree space
Published
• View Publication
• BIB
Phylogenetic networks are a generalization of phylogenetic trees that are used to represent reticulate evolution. Unrooted phylogenetic networks form a special class of such networks, which naturally generalize unrooted phylogenetic trees. In this paper we define two operations on unrooted phylogenetic networks, one of which is a generalization of the well-known nearest-neighbor interchange (NNI) operation on phylogenetic trees. We show that any unrooted phylogenetic network can be transformed into any other such network using only these operations. This generalizes the well-known fact that any phylogenetic tree can be transformed into any other such tree using only NNI operations. It also allows us to define a generalization of tree space and to define some new metrics on unrooted phylogenetic networks. To prove our main results, we employ some fascinating new connections between phylogenetic networks and cubic graphs that we have recently discovered. Our results should be useful in developing new strategies to search for optimal phylogenetic networks, a topic that has recently generated some interest in the literature, as well as for providing new ways to compare networks.
Bridging the gap between rooted and unrooted phylogenetic networks
The need for structures capable of accommodating complex evolutionary signals such as those found in, for example, wheat has fueled research into phylogenetic networks. Such structures generalize the standard phylogenetic tree model by also allowing cycles and have been introduced in rooted and unrooted form. In contrast to phylogenetic trees, however, surprisingly little is known about the interplay between both types thus hampering our ability to make much needed progress for rooted phylogenetic networks by drawing on insights from their much better understood unrooted counterparts. Unrooted phylogenetic networks are underpinned by split systems and by focusing on them we establish a first link between both types. More precisely, we develop a link between 1-nested phylogenetic networks which are examples of rooted phylogenetic networks and the well-studied median networks (aka Buneman graph) which are examples of unrooted phylogenetic networks. In particular, we show that not only can a 1-nested network be obtained from a median network but also that that network is, in a well-defined sense, optimal. Along the way, we characterize circular split systems in terms of the novel $\mathcal I$-intersection closure of a split system and establish the 1-nested analogue of the fundamental "Splits Equivalence Theorem" for phylogenetic trees.
Representing Partitions on Trees
Published
• View Publication
• BIB
In evolutionary biology, biologists often face the problem of constructing a phylogenetic tree on a set $X$ of species from a multiset $Π$ of partitions corresponding to various attributes of these species. One approach that is used to solve this problem is to try instead to associate a tree (or even a network) to the multiset $Σ_Π$ consisting of all those bipartitions $\{A,X-A\}$ with $A$ a part of some partition in $Π$. The rational behind this approach is that a phylogenetic tree with leaf set $X$ can be uniquely represented by the set of bipartitions of $X$ induced by its edges. Motivated by these considerations, given a multiset $Σ$ of bipartitions corresponding to a phylogenetic tree on $X$, in this paper we introduce and study the set $P(Σ)$ consisting of those multisets of partitions $Π$ of $X$ with $Σ_Π=Σ$. More specifically, we characterize when $P(Σ)$ is non-empty, and also identify some partitions in $P(Σ)$ that are of maximum and minimum size. We also show that it is NP-complete to decide when $P(Σ)$ is non-empty in case $Σ$ is an arbitrary multiset of bipartitions of $X$. Ultimately, we hope that by gaining a better understanding of the mapping that takes an arbitrary partition system $Π$ to the multiset $Σ_Π$, we will obtain new insights into the use of median networks and, more generally, split-networks to visualize sets of partitions.
Distinguished minimal toplogical lassos
Published
• View Publication
• BIB
A classical result in distance based tree-reconstruction characterizes when for a distance $D$ on some finite set $X$ there exist a uniquely determined dendrogram on $X$ (essentially a rooted tree $T=(V,E)$ with leaf set $X$ and no degree two vertices but possibly the root and an edge weighting $ω:E\to \mathbb R_{\geq 0}$) such that the distance $D_{(T,ω)}$ induced by $(T,ω)$ on $X$ is $D$. Moreover, algorithms that quickly reconstruct $(T,ω)$ from $D$ in this case are known. However in many areas where dendrograms are being constructed such as Computational Biology not all distances on $X$ are always available implying that the sought after dendrogram need not be uniquely determined anymore by the available distances with regards to topology of the underlying tree, edge-weighting, or both. To better understand the structural properties a set $\cL\subseteq {X\choose 2}$ has to satisfy to overcome this problem, various types of lassos have been introduced. Here, we focus on the question of when a lasso uniquely determines the topology of a dendrogram's underlying tree, that is, it is a topological lasso for that tree. We show that any set-inclusion minimal topological lasso for such a tree $T$ can be transformed into a 'distinguished' minimal topological lasso $\cL$ for $T$, that is, the graph $(X,\cL)$ is a claw-free block graph. Furthermore, we characterize such lassos in terms of the novel concept of a cluster marker map for $T$ and present results concerning the heritability of such lassos in the context of the subtree and supertree problems.
Reconstructing fully-resolved trees from triplet cover distances
Published
• View Publication
• BIB
It is a classical result that any finite tree with positively weighted edges, and without vertices of degree 2, is uniquely determined by the weighted path distance between each pair of leaves. Moreover, it is possible for a (small) strict subset $\cl$ of leaf pairs to suffice for reconstructing the tree and its edge weights, given just the distances between the leaf pairs in $\cl$. It is known that any set $\cl$ with this property for a tree in which all interior vertices have degree 3 must form a {\em cover} for $T$ -- that is, for each interior vertex $v$ of $T$, $\cl$ must contain a pair of leaves from each pair of the three components of $T-v$. Here we provide a partial converse of this result by showing that if a set $\cl$ of leaf pairs forms a cover of a certain type for such a tree $T$ then $T$ and its edge weights can be uniquely determined from the distances between the pairs of leaves in $\cl$. Moreover, there is a polynomial-time algorithm for achieving this reconstruction. The result establishes a special case of a recent question concerning `triplet covers', and is relevant to a problem arising in evolutionary genomics.
Lassoing and corraling rooted phylogenetic trees
Published in Bulletin of Mathematical Biology: Volume 75, Issue 3 (2013), Page 444-465
• View Publication
• BIB
The construction of a dendogram on a set of individuals is a key component of a genomewide association study. However even with modern sequencing technologies the distances on the individuals required for the construction of such a structure may not always be reliable making it tempting to exclude them from an analysis. This, in turn, results in an input set for dendogram construction that consists of only partial distance information which raises the following fundamental question. For what subset of its leaf set can we reconstruct uniquely the dendogram from the distances that it induces on that subset. By formalizing a dendogram in terms of an edge-weighted, rooted phylogenetic tree on a pre-given finite set X with |X|>2 whose edge-weighting is equidistant and a set of partial distances on X in terms of a set L of 2-subsets of X, we investigate this problem in terms of when such a tree is lassoed, that is, uniquely determined by the elements in L. For this we consider four different formalizations of the idea of "uniquely determining" giving rise to four distinct types of lassos. We present characterizations for all of them in terms of the child-edge graphs of the interior vertices of such a tree. Our characterizations imply in particular that in case the tree in question is binary then all four types of lasso must coincide.
Recognizing Treelike k-Dissimilarities
Published in Journal of Classification, 29 (2012), no. 3, 321-340
• View Publication
• BIB
A k-dissimilarity D on a finite set X, |X| >= k, is a map from the set of size k subsets of X to the real numbers. Such maps naturally arise from edge-weighted trees T with leaf-set X: Given a subset Y of X of size k, D(Y) is defined to be the total length of the smallest subtree of T with leaf-set Y . In case k = 2, it is well-known that 2-dissimilarities arising in this way can be characterized by the so-called "4-point condition". However, in case k > 2 Pachter and Speyer recently posed the following question: Given an arbitrary k-dissimilarity, how do we test whether this map comes from a tree? In this paper, we provide an answer to this question, showing that for k >= 3 a k-dissimilarity on a set X arises from a tree if and only if its restriction to every 2k-element subset of X arises from some tree, and that 2k is the least possible subset size to ensure that this is the case. As a corollary, we show that there exists a polynomial-time algorithm to determine when a k-dissimilarity arises from a tree. We also give a 6-point condition for determining when a 3-dissimilarity arises from a tree, that is similar to the aforementioned 4-point condition.
A Note on Encodings of Phylogenetic Networks of Bounded Level
Published
• View Publication
• BIB
Driven by the need for better models that allow one to shed light into the question how life's diversity has evolved, phylogenetic networks have now joined phylogenetic trees in the center of phylogenetics research. Like phylogenetic trees, such networks canonically induce collections of phylogenetic trees, clusters, and triplets, respectively. Thus it is not surprising that many network approaches aim to reconstruct a phylogenetic network from such collections. Related to the well-studied perfect phylogeny problem, the following question is of fundamental importance in this context: When does one of the above collections encode (i.e. uniquely describe) the network that induces it? In this note, we present a complete answer to this question for the special case of a level-1 (phylogenetic) network by characterizing those level-1 networks for which an encoding in terms of one (or equivalently all) of the above collections exists. Given that this type of network forms the first layer of the rich hierarchy of level-k networks, k a non-negative integer, it is natural to wonder whether our arguments could be extended to members of that hierarchy for higher values for k. By giving examples, we show that this is not the case.