arXiv++ Combinatorics

Browse math.CO papers from arXiv

phylogenetic

434 papers tagged with this keyword
2026-01-21 v2
A height-based metaconcept for rooted tree balance and its implications for the $B_1$ index
Tree balance has received considerable attention in recent years, both in phylogenetics and in other areas. Numerous (im)balance indices have been proposed to quantify the (im)balance of rooted trees. A recent comprehensive survey summarized this literature and showed that many existing indices are based on similar underlying principles. To unify these approaches, three general metaconcepts were introduced, providing a framework to classify, analyze, and extend imbalance indices. In this context, a metaconcept is a function $Φ_f$ that depends on another function $f$ capturing some aspect of tree shape. In this manuscript, we extend this line of research by introducing a new metaconcept based on the heights of the pending subtrees of all inner vertices. We provide a thorough analysis of this metaconcept and use it to answer open questions concerning the well-known $B_1$ balance index. In particular, we characterize the tree shapes that maximize the $B_1$ index in two cases: (i) arbitrary rooted trees and (ii) binary rooted trees. For both cases, we also determine the corresponding maximum values of the index. Finally, while the $B_1$ index is induced by a so-called third-order metaconcept, we explicitly introduce three new (im)balance indices derived from the first- and second-order height metaconcepts, respectively, thereby demonstrating that pending subtree heights give rise to a variety of novel (im)balance indices.
2026-01-14 v2
Proof of a Conjecture on Young Tableaux with Walls
Banderier, Marchal, and Wallner considered Young tableaux with walls, which are similar to standard Young tableaux, except that local decreases are allowed at some walls. In this work, we prove a conjecture of Fuchs and Yu concerning the enumeration of two classes of three-row Young tableaux with walls. Combining with the work by Chang, Fuchs, Liu, Wallner, and Yu leads to the verification of a conjecture on tree-child networks proposed by Pons and Batle. This conjecture was regarded as a specific and challenging problem in the Phylogenetics community until it was finally resolved by the present work.
2026-01-13
Metric properties of electrical networks and the graph reconstruction problems
Published in In EuroComb'25, Booklet of extended abstracts, pages 456-460 . HUN-REN Alfréd Rényi Institute of Mathematics, 2025 • Search Publication
Using the generalized Temperley trick, we demonstrate the explicit embedding of circular electrical networks into totally non-negative Grassmannians. Building on this result, we show that the effective resistances between boundary nodes of circular electrical networks satisfy the Kalmanson property, and we provide the full characterization of planar electrical Kalmanson metrics. Additionally, we present a graph reconstruction algorithm with applications in phylogenetic network analysis as well as the numerical solution of the Calderon problem.
Combinatorial comparison of general galled trees, time-consistent galled trees, and simplex time-consistent galled trees
Published • View Publication • BIB
Rooted binary phylogenetic networks are extensions of rooted binary trees, adding reticulation nodes that are designed to represent evolutionary processes that involve hybridization events. Enumerative combinatorics studies have counted leaf-labeled phylogenetic networks in a variety of classes, finding that when the number of reticulations is fixed, the time-consistent galled trees are asymptotically less numerous than each of several network classes that had been previously examined. Here we provide enumerative results on two additional network classes: general galled trees and simplex time-consistent galled trees. We show that for a fixed number of galls, as the number of leaves goes to infinity, the asymptotic count of general galled trees is identical to that of time-consistent galled trees, whereas the count of simplex time-consistent galled trees is smaller. If the number of galls is not restricted, then the asymptotic approximations all differ: simplex time-consistent galled trees are less numerous than time-consistent galled trees, which are in turn less numerous than general galled trees. We also report a variety of additional results: recursions to count the studied networks with small numbers of leaves a fixed number of galls, as well as enumerative results for unlabeled networks in the classes that we investigate.
Arboreal Ultrametrics
Ultametrics are an important class of distances used in applications such as phylogenetics, clustering and classification theory. Ultrametrics are essentially distances that can be represented by an edge-weighted rooted tree so that all of the distances in the tree from the root to any leaf of the tree are equal. In this paper, we introduce a generalization of ultrametrics called arboreal ultrametrics which have applications in phylogenetics and also arise in the theory of distance-hereditary graphs. These are partial distances, that is distances that are not necessarily defined for every pair of elements in the groundset, that can be represented by an ultrametric arboreal network, that is, an edge-weighted rooted network whose underlying graph is a tree. As with ultrametrics all of the distances in the ultrametric arboreal network from any root to any leaf below it are are equal but, in contrast, the network may have more than one root. In our two main results we characterize when a partial distance is an arboreal ultrametric as well as proving that, somewhat surprisingly, given any unrooted edge-weighted phylogenetic tree there is a necessarily unique way to insert roots into this tree so as to obtain an arboreal ultrametric.
2025-12-15
Binary normal networks without near reticulations can be reconstructed from their rooted triples
Normal networks are an important class of phylogenetic networks that have compelling mathematical properties which align with intuition about inference from genetic data. While tools enabling widespread use of phylogenetic networks in the biological literature are still under mathematical, statistical, and computational development, many such results are being assembled, and in particular for normal phylogenetic networks. For instance, it has been shown that binary normal networks can be reconstructed from the sets of three- and four-leaf rooted phylogenetic trees that they display. It is also known that one can reconstruct particular subclasses of normal networks from just the displayed rooted triples. This applies, for instance, to rooted binary phylogenetic trees and to binary level-$1$ normal networks. In this paper we address the question of how much of the class of binary normal networks can be reconstructed from just the rooted triples that they display. We find that all except those with substructures that we call ``near-sibling reticulations'' and ``near-stack reticulations'' can be reconstructed just from their rooted triples. This goes some way to answering the natural question of how much information can be extracted from a set of displayed rooted triples, which are arguably the simplest substructure that one may hope for in a phylogenetic object.
Bounds on the sequence length sufficient to reconstruct level-1 phylogenetic networks
Published • View Publication • BIB
Phylogenetic trees and networks are graphs used to model evolutionary relationships, with trees representing strictly branching histories and networks allowing for events in which lineages merge, called reticulation events. While the question of data sufficiency has been studied extensively in the context of trees, it remains largely unexplored for networks. In this work we take a first step in this direction by establishing bounds on the amount of genomic data required to reconstruct binary level-$1$ semi-directed phylogenetic networks, which are binary networks in which reticulation events are indicated by directed edges, all other edges are undirected, and cycles are vertex-disjoint. For this class, methods have been developed recently that are statistically consistent. Roughly speaking, such methods are guaranteed to reconstruct the correct network assuming infinitely long genomic sequences. Here we consider the question whether networks from this class can be uniquely and correctly reconstructed from finite sequences. Specifically, we present an inference algorithm that takes as input genetic sequence data, and demonstrate that the sequence length sufficient to reconstruct the correct network with high probability, under the Cavender-Farris-Neyman model of evolution, scales logarithmically, polynomially, or polylogarithmically with the number of taxa, depending on the parameter regime. As part of our contribution, we also present novel inference rules for quartet data in the semi-directed phylogenetic network setting.
2025-11-20
Labeled histories and maximally probable labeled topologies with multifurcation
Published • View Publication • BIB
In mathematical phylogenetics, labeled histories describe the sequences by which sets of labeled lineages coalesce to a shared ancestral lineage. We study labeled histories for at-most-$r$-furcating trees. Consider a rooted leaf-labeled tree in which internal nodes each have $i$ offspring, and $i$ is permitted to range from 2 to $r$ across internal nodes, for a specified value of $r$. For labeled topologies with $n$ leaves, we enumerate the total number of labeled histories with at-most-$r$-furcation. We enumerate the labeled histories possessed by a specific at-most-$r$-furcating labeled topology. We then demonstrate that the maximally probable at-most-$r$-furcating unlabeled topology on $n \geq 2$ leaves -- the unlabeled topology whose labelings have the largest number of labeled histories -- is the maximally probable strictly bifurcating unlabeled topology on $n$ leaves. Finally, we enumerate labeled histories for at-most-$r$-furcating labeled topologies in a setting that permits simultaneous branchings. We similarly reduce the problem of identifying the maximally probable at-most-$r$-furcating unlabeled topology on $n \geq 2$ leaves, allowing simultaneity, to that of identifying the maximally probable strictly bifurcating unlabeled topology on $n$ leaves, with simultaneity; we conjecture the shape of this bifurcating unlabeled topology. The computations contribute to the study of multifurcation, which arises in various biological processes, and they connect to analogous mathematical settings involving precedence-constrained scheduling.
Inferring DAGs and Phylogenetic Networks from Least Common Ancestors
Published • View Publication • BIB
A least common ancestor (LCA) of two leaves in a directed acyclic graph (DAG) is a vertex that is an ancestor of both leaves and has no proper descendant that is also their common ancestor. LCAs capture hierarchical relationships in rooted trees and, more generally, in DAGs. In 1981, Aho et al. introduced the problem of determining whether a set of pairwise LCA constraints on a set $X$, of the form $(i,j)<(k,l)$ with $i,j,k,l\in X$, can be realized by a rooted tree whose leaf set is $X$, such that whenever $(i,j)<(k,l)$, the LCA of $i,j$ is a descendant of that of $k,l$. They also presented a polynomial-time algorithm, BUILD, to solve this problem. However, many such constraint systems cannot be realized by any tree, prompting the question of whether they can be realized by a more general DAG. We extend Aho et al.'s framework from trees to DAGs, providing both theoretical and algorithmic foundations for reasoning about LCA constraints in this broader setting. Given a collection $R$ of LCA constraints, we define its $+$-closure $R^+$, capturing additional LCA relations implied by $R$. Using $R^+$, we construct a canonical DAG $G_R$ and prove that $R$ is DAG-realizable if and only if it is realized by $G_R$. We further adapt this construction to phylogenetic networks, defining a canonical network $N_R$ and prove that it is regular, i.e., it coincides with the Hasse diagram of its underlying set system. Finally, we show that for any DAG-realizable $R$, its classical closure - comprising all LCA constraints that hold in every DAG realizing $R$ - coincides with its $+$-closure. All constructions are computable in polynomial time, and we provide explicit algorithms for each.
Characterizations of undirected 2-quasi best match graphs
Bipartite best match graphs (BMG) and their generalizations arise in mathematical phylogenetics as combinatorial models describing evolutionary relationships among related genes in a pair of species. In this work, we characterize the class of \emph{undirected 2-quasi-BMGs} (un2qBMGs), which form a proper subclass of the $P_6$-free chordal bipartite graphs. We show that un2qBMGs are exactly the class of bipartite graphs free of $P_6$, $C_6$, and the eight-vertex Sunlet$_4$ graph. Equivalently, a bipartite graph $G$ is un2qBMG if and only if every connected induced subgraph contains a ``heart-vertex'' which is adjacent to all the vertices of the opposite color. We further provide a $O(|V(G)|^3)$ algorithm for the recognition of un2qBMGs that, in the affirmative case, constructs a labeled rooted tree that ``explains'' $G$. Finally, since un2qBMGs coincide with the $(P_6,C_6)$-free bi-cographs, they can also be recognized in linear time.
Generalizing matrix representations to fully heterochronous ranked tree shapes
Published • View Publication • BIB
Phylogenetic tree shapes capture fundamental signatures of evolution. We consider ``ranked'' tree shapes, which are equipped with a total order on the internal nodes compatible with the tree graph. Recent work has established an elegant bijection of ranked tree shapes and a class of integer matrices, called \textbf{F}-matrices, defined by simple inequalities. This formulation is for isochronous ranked tree shapes, where all leaves share the same sampling time, such as in the study of ancient human demography from present-day individuals. Another important style of phylogenetics concerns trees where the ``timing'' of events is by branch length rather than calendar time. This style of tree, called a rooted phylogram, is output by popular maximum-likelihood methods. These trees are broadly relevant, such as to study the affinity maturation of B cells in the immune system. Discretizing time in a rooted phylogram gives a fully heterochronous ranked tree shape, where leaves are part of the total order. Here we extend the \textbf{F}-matrix framework to such fully heterochronous ranked tree shapes. We establish an explicit bijection between a class of \textbf{F}-matrices and the space of such tree shapes. The matrix representation has the key feature that values at any entry are highly constrained via four previous entries, enabling straightforward enumeration of all valid tree shapes. We also use this framework to develop probabilistic models on ranked tree shapes. Our work extends understanding of combinatorial objects that have a rich history in the literature.
2025-10-23 v2
Labeling and folding multi-labeled trees
In 1989 Erdős and Székely showed that there is a bijection between (i) the set of rooted trees with $n+1$ vertices whose leaves are bijectively labeled with the elements of $[\ell]=\{1,2,\dots,\ell\}$ for some $\ell \leq n$, and (ii) the set of partitions of $[n]=\{1,2,\dots,n\}$. They established this via a labeling algorithm based on the anti-lexicographic ordering of non-empty subsets of $[n]$ which extends the labeling of the leaves of a given tree to a labeling of all of the vertices of that tree. In this paper, we generalize their approach by developing a labeling algorithm for multi-labeled trees, that is, rooted trees whose leaves are labeled by positive integers but in which distinct leaves may have the same label. In particular, we show that certain orderings of the set of all finite, non-empty multisets of positive integers can be used to characterize partitions of a multiset that arise from labelings of multi-labeled trees. As an application, we show that the recently introduced class of labelable phylogenetic networks is precisely the class of phylogenetic networks that are stable relative to the so-called folding process on multi-labeled trees. We also give a bijection between the labelable phylogenetic networks with leaf-set $[n]$ and certain partitions of multisets.
Recognizing Leaf Powers and Pairwise Compatibility Graphs is NP-Complete
Published • View Publication • BIB
Leaf powers and pairwise compatibility graphs were introduced over twenty years ago as simplified graph models for phylogenetic trees. Despite significant research, several properties of these graph classes remain poorly understood. In this paper, we establish that the recognition problem for both classes is NP-complete. We extend this hardness result to a broader hierarchy of graph classes, including pairwise compatibility graphs and their generalizations, multi interval pairwise compatibility graphs.
2025-10-10 v2
Parameterized Algorithms for Diversity of Networks with Ecological Dependencies
For a phylogenetic tree, the phylogenetic diversity of a set A of taxa is the total weight of edges on paths to A. Finding small sets of maximal diversity is crucial for conservation planning, as it indicates where limited resources can be invested most efficiently. In recent years, efficient algorithms have been developed to find sets of taxa that maximize phylogenetic diversity either in a phylogenetic network or in a phylogenetic tree subject to ecological constraints, such as a food web. However, these aspects have mostly been studied independently. Since both factors are biologically important, it seems natural to consider them together. In this paper, we introduce decision problems where, given a phylogenetic network, a food web, and integers k, and D, the task is to find a set of k taxa with phylogenetic diversity of at least D under the maximize all paths measure, while also satisfying viability conditions within the food web. Here, we consider different definitions of viability, which all demand that a "sufficient" number of prey species survive to support surviving predators. We investigate the parameterized complexity of these problems and present several fixed-parameter tractable (FPT) algorithms. Specifically, we provide a complete complexity dichotomy characterizing which combinations of parameters - out of the size constraint k, the acceptable diversity loss D, the scanwidth of the food web, the maximum in-degree in the network, and the network height h - lead to W[1]-hardness and which admit FPT algorithms. Our primary methodological contribution is a novel algorithmic framework for solving phylogenetic diversity problems in networks where dependencies (such as those from a food web) impose an order, using a color coding approach.
2025-09-19
Ordered Leaf Attachment (OLA) Vectors can Identify Reticulation Events even in Multifurcated Trees
Recently, a new vector encoding, Ordered Leaf Attachment (OLA), was introduced that represents $n$-leaf phylogenetic trees as $n-1$ length integer vectors by recording the placement location of each leaf. Both encoding and decoding of trees run in linear time and depend on a fixed ordering of the leaves. Here, we investigate the connection between OLA vectors and the maximum acyclic agreement forest (MAAF) problem. A MAAF represents an optimal breakdown of $k$ trees into reticulation-free subtrees, with the roots of these subtrees representing reticulation events. We introduce a corrected OLA distance index over OLA vectors of $k$ trees, which is easily computable in linear time. We prove that the corrected OLA distance corresponds to the size of a MAAF, given an optimal leaf ordering that minimizes that distance. Additionally, a MAAF can be easily reconstructed from optimal OLA vectors. We expand these results to multifurcated trees: we introduce an $O(kn \cdot m\log m)$ algorithm that optimally resolves a set of multifurcated trees given a leaf-ordering, where $m$ is the size of a largest multifurcation, and show that trees resolved via this algorithm also minimize the size of a MAAF. These results suggest a new approach to fast computation of phylogenetic networks and identification of reticulation events via random permutations of leaves. Additionally, in the case of microbial evolution, a natural ordering of leaves is often given by the sample collection date, which means that under mild assumptions, reticulation events can be identified in polynomial time on such datasets.
2025-09-05
Revealing the building blocks of tree balance: fundamental units of the Sackin and Colless Indices
Over the past decades, more than 25 phylogenetic tree balance indices and several families of such indices have been proposed in the literature -- some of which even contain infinitely many members. It is well established that different indices have different strengths and perform unequally across application scenarios. For example, power analyses have shown that the ability of an index to detect the generative model of a given phylogenetic tree varies significantly between indices. This variation in performance motivates the ongoing search for new and possibly \enquote{better} (im)balance indices. An easy way to generate a new index is to construct a compound index, e.g., a linear combination of established indices. Two of the most prominent and widely used imbalance indices in this context are the Sackin index and the Colless index. In this study, we show that these classic indices are themselves compound in nature: they can be decomposed into more elementary components that independently satisfy the defining properties of a tree (im)balance index. We further show that the difference Colless minus Sackin results in another imbalance index that is minimized (amongst others) by all Colless minimal trees. Conversely, the difference Sackin minus Colless forms a balance index. Finally, we compare the building blocks of which the Sackin and the Colless indices consist to these indices as well as to the stairs2 index, which is another index from the literature. Our results suggest that the elementary building blocks we identify are not only foundational to established indices but also valuable tools for analyzing disagreement among indices when comparing the balance of different trees.
2025-08-29 v2
When Many Trees Go to War: On Sets of Phylogenetic Trees With Almost No Common Structure
Published • View Publication • BIB
It is known that any two trees on the same $n$ leaves can be displayed by a network with $n-2$ reticulations, and there are two trees that cannot be displayed by a network with fewer reticulations. But how many reticulations are needed to display multiple trees? For any set of $t$ trees on $n$ leaves, there is a trivial network with $(t - 1)n$ reticulations that displays them. To do better, we have to exploit common structure of the trees to embed non-trivial subtrees of different trees into the same part of the network. In this paper, we show that for $t \in o(\sqrt{\lg n})$, there is a set of $t$ trees with virtually no common structure that could be exploited. More precisely, we show for any $t\in o(\sqrt{\lg n})$, there are $t$ trees such that any network displaying them has $(t-1)n - o(n)$ reticulations. For $t \in o(\lg n)$, we obtain a slightly weaker bound. We also prove that already for $t = c\lg n$, for any constant $c > 0$, there is a set of $t$ trees that cannot be displayed by a network with $o(n \lg n)$ reticulations, matching up to constant factors the known upper bound of $O(n \lg n)$ reticulations sufficient to display \emph{all} trees with $n$ leaves. These results are based on simple counting arguments and extend to unrooted networks and trees.
2025-08-21
Defining a phylogenetic tree with the minimum number of small-state characters
Phylogenetic trees represent evolutionary relationships and can be uniquely defined by sets of finite-state biological characteristics. Despite prior work showing that sufficiently large trees can be determined by $r$-state character sets, the minimal leaf thresholds $n_r$ remain largely unknown. In this work, we establish the 3-state case as $n_3 = 8$, providing a concrete base for higher-state analyses. We then resolve the 5-state problem by constructing a counterexample for $n=15$ and proving that for $n \geq 16$, $\lceil (n-3)/4 \rceil$ 5-state characters suffice to uniquely define any tree. Our approach relies on rigorous mathematical induction with complete verification of base cases and logically consistent inductive steps, offering new insights into the minimal conditions for character-based tree identification.
2025-08-19
A sharp lower bound for the number of phylogenetic trees displayed by a tree-child network
Published • View Publication • BIB
A normal (phylogenetic) network with $k$ reticulations displays $2^k$ phylogenetic trees. In this paper, we establish an analogous result for tree-child (phylogenetic) networks with no underlying $3$-cycles. In particular, we show that a tree-child network with $k\ge 2$ reticulations and no underlying $3$-cycles displays at least $2^{k/2}$ phylogenetic trees if $k$ is even and at least $\frac{3}{2\sqrt{2}}2^{k/2}$ if $k$ is odd. Moreover, we show that these bounds are sharp and characterise the tree-child networks that attain these bounds.
2025-08-07
Identifiability of Large Phylogenetic Mixtures for Many Phylogenetic Model Structures
Identifiability of phylogenetic models is a necessary condition to ensure that the model parameters can be uniquely determined from data. Mixture models are phylogenetic models where the probability distributions in the model are convex combinations of distributions in simpler phylogenetic models. Mixture models are used to model heterogeneity in the substitution process in DNA sequences. While many basic phylogenetic models are known to be identifiable, mixture models in generality have only been shown to be identifiable in certain cases. We expand the main theorem of [Rhodes, Sullivant 2012] to prove identifiability of mixture models in equivariant phylogenetic models, specifically the Jukes-Cantor, Kimura 2-parameter model, Kimura 3-parameter model and the Strand Symmetric model.