arXiv++ Combinatorics

Browse math.CO papers from arXiv

phylogenetic

434 papers tagged with this keyword
Encoding level-3 semi-directed phylogenetic networks by quarnets and quinnets
Phylogenetic networks generalize phylogenetic trees as models of evolutionary history, allowing lineages to merge as well as to diverge. For many types of genetic data the root position of such a network cannot be recovered, so that only a semi-directed network can be inferred: a mixed graph in which only the edges entering a reticulation vertex are directed. A common strategy for inferring such a network is to first infer the subnetwork it induces on each set of $k\geq 3$ of its leaves, called a $k$-net, and then to assemble these pieces. This can only succeed if the $k$-nets determine the network, in which case that network is said to be encoded by its $k$-nets. Semi-directed networks of level-1 and 2, those whose biconnected components contain at most one, respectively two, reticulations, are known to be encoded by their $4$-nets, or quarnets, whereas level-3 networks are not. Even so, in this paper we show that level-3 semi-directed networks are encoded by their $5$-nets, or quinnets, and we characterize the limitation of quarnets exactly: we show that a single previously reported counterexample captures the only obstruction, every other level-3 network being encoded by its quarnets. Our proofs rest on a collection of encoding results for individual structural features of a network, which we establish for networks of arbitrary level and which are of independent interest.
2026-10-02
Exact and asymptotic enumeration of unrestricted binary phylogenetic networks through automorphism weights
Let $\cP_{\ell,k}$ be the set of rooted binary phylogenetic networks with $\ell$ labelled leaves, $k$ reticulations and no parallel edges. We write $|\cP_{\ell,k}|=W_k(\ell)+D_k(\ell)$, where the weighted count $W_k(\ell)$ adds the inverse orders of the leaf-fixing automorphism groups and the defect $D_k(\ell)$ collects the remainder. The weighted count satisfies, for every $k$, a recursion over the source layers of the tree-component structure that involves neither a list of component graphs nor any distinction between symmetric and asymmetric configurations, and its exponential generating function is a Laurent polynomial in $\sqrt{1-2x}$. Automorphism groups of networks are $2$-groups, elementary abelian for $k\le5$ but not in general. For $k\le5$ the defect is the weighted count of networks with a distinguished involution, which obeys an extension of the same recursion. An exact symbolic evaluation of the two recursions yields $|\cP_{\ell,k}|$ in closed form for $k\le5$, the case $k=5$ being new; it reproduces the published counts for $k\le4$ and corrects a coefficient in a published generating function for $k=3$. For every $k$, uniformly over explicit ranges of $k$, we prove that the non-tree-child networks are a fraction $2k(k-1)/\ell$ of the tree-child networks to leading order, which gives the third term of the asymptotic expansion of $|\cP_{\ell,k}|$. We also prove that reticulation-visible networks exceed tree-child networks by the fraction $k(k-1)/\ell$, and that a uniformly random network in $\cP_{\ell,k}$ has a nontrivial automorphism with probability $k(k-1)/(4\ell^3)$ to leading order.
2026-10-02 v3
Exact Enumeration of Phylogenetic Networks: The Tree-Child, Reticulation-Visible and Orchard Hierarchy
We develop a unified framework for the exact enumeration and asymptotic analysis of the three most studied classes of phylogenetic networks: tree-child (TC), reticulation-visible (RV) and orchard networks, whose cardinalities satisfy the strict ordering $|\mathrm{TC}_{\ell,k}|<|\mathrm{RV}_{\ell,k}|<|\mathrm{Orch}_{\ell,k}|$ for reticulation number $k\geq2$ (with $\mathrm{TC}\subsetneq\mathrm{RV}$ and $\mathrm{TC}\subsetneq\mathrm{Orch}$, while $\mathrm{RV}$ and $\mathrm{Orch}$ are incomparable as sets). Using the Chang--Fuchs structural theorem, we derive a two-level master functional equation for the RV bivariate generating function and obtain exact closed-form identities for the differences $Δ_k(\ell):=|RV_{\ell,k}|-|TC_{\ell,k}|$ for $k=2,3$, with the asymptotic universality $Δ_k(\ell)/|TC_{\ell,k}|\sim k!/\ell$. For orchard networks, we prove a \emph{universal hypergeometric law} that resolves the exact enumeration problem for all $\ell$: the column generating function $F_\ell(v)$ is rational with denominator $D_\ell(v)=\prod_{j=2}^\ell X_j(v)$, where \[ X_\ell(v) = \sum_{k=0}^{\lfloor\ell/2\rfloor}(-1)^k\, \frac{\ell!}{(\ell-2k)!\,k!}\,v^k \] is the matching polynomial of the complete graph $K_\ell$ and a rescaled Jacobi polynomial. This immediately resolves the intractable $\ell=9$ case: $D_9$ has degree 20, dominant growth rate $\approx40.73$, and all spectral roots are positive real. A complete enumeration table is provided extending the published data of Cardona, Ribas and Pons.
2026-10-01 v2
Characterisations of Planar Galled Networks
Rooted phylogenetic networks are widely used to represent the evolution of species that have undergone reticulate processes. However, these networks can be highly non-planar, making them more difficult to visualise and interpret than evolutionary trees. In this paper, we investigate planarity properties of galled networks, an important subclass of phylogenetic networks. We show that all planar galled networks are necessarily upward planar. Furthermore, by leveraging recent results on planar phylogenetic networks, we provide three characterisations for each of the outerplanar and terminal planar galled network classes in terms of forbidden vertex configurations, forbidden directed subgraphs, and forbidden structures in their associated underlying undirected graphs. These results contribute to a deeper understanding of the structural properties of galled networks and may inform future methods for their construction and visualisation.
2026-10-01 v2
Exact Counts of Binary Phylogenetic Networks with Four Reticulations
Phylogenetic networks provide a flexible framework for representing reticulate evolutionary processes, such as hybridization, introgression, recombination, and horizontal gene transfer. However, their combinatorial complexity makes even basic enumeration problems difficult. Building on our previous work for networks with up to three reticulations, we derive an explicit closed-form formula for the number of unrestricted rooted binary phylogenetic networks with four reticulations on \(n\) labeled taxa. Our approach is based on tree-component graphs. We classify the 79 possible component graphs corresponding to networks with four reticulations into ten groups. We then enumerate the networks associated with each group by combining known counts of one-component networks, forests, and networks with fewer reticulations. Summing these contributions yields the desired formula. This result extends the exact enumeration of unrestricted binary phylogenetic networks to four reticulations and further demonstrates the effectiveness of component graphs for systematically organizing and counting increasingly complex network classes.
Characterization of tree-child networks in terms of mu vectors
We characterize tree-child phylogenetic networks in terms of their mu-representations. First, we give a structural characterization of tree-child networks by means of ordered tree-path decompositions. We then translate this decomposition into a set of purely vectorial conditions on finite subsets M in N^n. We prove that such a set M is the mu-representation of a tree-child phylogenetic network if and only if it is tree-child mu-compatible. This provides a feasibility criterion for tree-child mu-representations which can be used as a basis for reconstruction and further algorithmic applications. Note that this paper presents results arising from ongoing research on tree-child networks and that the results will be further developed and placed into proper context in subsequent versions.
2026-09-21 v3
Asymptotics for the harmonic descent chain and applications to critical beta-splitting trees
Motivated by the connection to a probabilistic model of phylogenetic trees introduced by Aldous, we study the recursive sequence governed by the rule $x_n = \sum_{i=1}^{n-1} \frac{1}{h_{n-1}(n-i)} x_i$ where $h_{n-1} = \sum_{j=1}^{n-1} 1/j$, known as the harmonic descent chain. While it is known that this sequence converges to an explicit limit $x$, not much is known about the rate of convergence. We first show that a class of recursive sequences including the above are decreasing and use this to bound the rate of convergence. Moreover, for the harmonic descent chain we prove the asymptotic $x_n - x = n^{-γ_* + o(1)}$ for an implicit exponent $γ_*$. As a consequence, we deduce central limit theorems for various statistics of the critical beta-splitting random tree. This answers a number of questions of Aldous, Janson, and Pittel.
The phylogenetic rank of a graph
The Pachter-Sturmfels phylogenetic rank of a graph G is the minimal number of metric trees needed to embed G isometrically. Here, all edges of G are of length one and the product of metric trees is endowed with the supremum norm. We develop both a greedy and an exact algorithm for computing phylogenetic ranks. Using our algorithms, we construct a database of phylogenetic ranks which includes all graphs on 6 and 7 vertices. In particular, we exhibit examples disproving that the phylogenetic rank is hereditary, bounded by $\lceil \frac{n}{2} \rceil$, and a generalised 4-point conjecture by Pachter and Sturmfels. In addition, we show that the phylogenetic rank is subadditive under 1-sums and certain 2-vertex-sums, and that it is trivially upper bounded by n-1, where n denotes the number of vertices. We also provide a complete classification of graphs with phylogenetic rank 1 and construct several infinite families with phylogenetic rank $\lceil \frac{n}{2} \rceil$.
2026-09-11
NOC NOC, who's there? Clustering systems of tree-child and normal networks
Clustering systems provide a natural way to encode structural information contained in phylogenetic networks. In this note, we study the clustering systems of normal and tree-child networks through an overlap-based property of set systems, called not-overlap-covered (NOC). We show that the NOC property is equivalent to inclusion-visibility, a memberwise formulation of the strict-compatibility condition previously used for tree-child clustering systems. We characterize normal networks as precisely the semi-regular networks whose clustering systems satisfy NOC. Consequently, a clustering system is realized by a normal network if and only if it satisfies NOC, or equivalently, if every one of its clusters is inclusion-visible. In this case, the Hasse diagram provides a canonical normal realization. These are exactly the clustering systems realized by tree-child networks. The NOC formulation yields a sharp quadratic upper bound on the number of distinct clusters of tree-child and normal networks and a direct polynomial-time recognition algorithm. Finally, we explore several consequences of the NOC perspective beyond the phylogenetic setting. These include connections to the enumeration of normal networks, an order-theoretic interpretation of inclusion-visibility, structural properties of NOC set systems, and a tractable special case of Minimum Set Cover, which is NP-hard in general.
2026-09-09
Hyperbolic distance matrix completion
A completion theory for hyperbolic distance data is developed at the interface of matrix analysis, graph theory, and hyperbolic geometry. Krein's characterization of the metric space embeddability in Lobachevsky space leads to a natural anchoring procedure that transforms the indefinite data into a positive semidefinite kernel. In analogy with positive semidefinite and Euclidean distance matrix completion, chordality of the specification graph is shown to be the necessary and sufficient condition for local Lorentz-Gram data to admit global completion. Existence is complemented by explicit constructions. For trees, we obtain geodesic-rectification and product-distance completions; for chordal graphs, the latter extends to matrix-valued transfers along clique-trees. The resulting canonical completion is characterized by sparsity of its inverse and by a maximum-absolute-determinant principle. Its metric distortion exhibits a sharp dichotomy governed by clique separator size. Applications to exact recovery from sparse hyperbolic measurements and to hierarchical and phylogenetic data are developed.
2026-09-09 v2
Regularizing and Normalizing DAGs and Phylogenetic Networks
Phylogenetic networks and, more generally, directed acyclic graphs (DAGs) represent hierarchical structure beyond trees, for instance in the presence of reticulate evolutionary events such as hybridization or horizontal gene transfer. A central question is which parts of such graphs are essential with respect to leaf-observable information, and which parts can be removed without changing this information. Resolving this question can lead to principled simplification methods for phylogenetic networks, such as the recent normalization approach of Francis et al. In this paper, we study this question from three related perspectives: clusters displayed by a DAG $G$, least common ancestors (LCAs) of subsets of its leaf set, and visibility, a path-based property of vertices. We first introduce an LCA-based simplification procedure called $i$-regularization. For a DAG $G$ and $i\geq 1$, the DAG $\reg_i(G)$ retains precisely those vertices that occur as unique LCAs of leaf subsets of size at most $i$, removes the remaining non-leaf vertices by a graph-editing operation $\ominus$, and then deletes shortcuts. We show that $\reg_i(G)$ admits a Hasse-diagram characterization in terms of the corresponding lca-clusters. We then compare LCA-based regularization with normalization. Using the same $\ominus$-operator, we describe the cover construction underlying normalization, identify visible vertices that are nevertheless removed, and characterize when regularization and normalization coincide. Together, these results provide a unified framework for cluster-based, LCA-based, and visibility-based simplifications of DAGs and phylogenetic networks.
2026-09-04
A Short Combinatorial Proof of the Pons-Batle Identity for Counting Tree-Child Networks
Tree-child networks are a useful class of binary phylogenetic networks. The Pons--Batle identity (Pons and Batle, \textit{Scientific Reports}, 2021) states that the number $a_{n,k}$ of tree-child networks with $k$ reticulations on $n$ taxa satisfies \[ a_{n,k}=(n-k+1)a_{n,k-1} +\frac{n(2n+k-3)}{n-k}a_{n-1,k}. \] In this paper, we present a short combinatorial proof of this identity.
2026-09-02
An Affine Semigroup from Orbifold Boundary Conditions: cut, phylogenetic and hierarchical models in the unit-weight sector, and weighted configurations beyond them
The equivalence classes of boundary conditions of a gauge theory on a two-dimensional orbifold are the fibres of a marginal map, indexed by an affine semigroup: one generator per alphabet label, graded by weight, embedded by its local data at the fixed points. This note identifies that semigroup. Without weights the configuration has a name and a literature, whose results about our cases are attributed here: over $\mathbb{Z}_2$ it is the cut configuration of an explicit graph in the sense of Sturmfels-Sullivant --- the four-cycle for $T^2/\mathbb{Z}_2$, the wheel $W_4$ for $S^1/\mathbb{Z}_2\times S^1/\mathbb{Z}_2$ --- verified as an equality of configurations; over $\mathbb{Z}_m$ with equal cone orders, the group-based phylogenetic model on a claw tree; with unequal orders, a mixed-order variant we do not find in the literature; for higher products, the binary hierarchical model of a cross-polytope boundary complex. The product orbifold's ring is a row of a 2008 table --- codimension, degree, minimal generators, normality --- every invariant of which our machinery reproduced without knowing of it. What none of the three covers is the alphabet with weights, which arise from induction to higher-dimensional irreducibles of a non-abelian space group and from conjugate-pair recombination over real or quaternionic ground. That sector is adjacent to, but not identified with, the non-abelian direction Sturmfels and Sullivant raised in 2005, and is where our contributions sit: gluing trees for the weighted alphabets and the orthogonal and symplectic columns, and the group-based model on the tripod, a complete intersection exactly when the finite abelian group has order at most three. The first group beyond $\mathbb{Z}_3$ separates local from global: the $\mathbb{Z}_4$ tripod is a complete intersection on the Zariski-open set the phylogenetics literature works in, and not globally.
2026-08-26
Tree Buckets and the Reconstruction of Pairs of Phylogenetic Trees
Phylogenetic trees are used in evolutionary biology to represent the evolutionary history of a collection of taxa. As we have incomplete information about any evolutionary history, recovering trees from partial information is a focus of phylogenetic combinatorics. However, in some cases the available data does not describe a single phylogenetic tree. We consider recovery of pairs of phylogenetic trees from their combined subtrees with $k$ leaves, which we call a $k$-bucket. We establish the exact cases in which these pairs of trees are recoverable from their subtrees with a single leaf removed, both when just considering the structure of the trees, and when additionally considering the set of taxa on the leaves. We also consider recovery of pairs of trees with labelled leaves from their rooted triples, and establish that they are recoverable up to a sequence of subtree swaps.
2026-08-24
On the maximum size of 2-weakly compatible split systems
We consider a Turán-type problem arising in phylogenetics: determining the maximum size of a 2-weakly compatible split system. This compatibility condition arises in the reconstruction of phylogenetic networks from quartet weights. It was previously shown that a 2-weakly compatible split system has size at most \[ 3\binom{n}{4}+\binom{n}{2}. \] We prove that the maximum size is $O(n^{5/2})$.
2026-08-21
A cube-root phase transition in tree-child networks and the enumeration threshold for galled networks
We prove two surprising results about phylogenetic networks. First, we show that the structure of tree-child networks with $n$ leaves and $k$ reticulation nodes undergoes a sharp phase transition at $n^{1/3}$: if $k=o(n^{1/3})$, then a random tree-child network is almost surely a semi-simplex tree-child network, whereas if $k/n^{1/3}\rightarrow\infty$ and $k=o(n^{1/2})$, it is almost surely not. Second, we show that this result implies that the asymptotic counting formula for galled networks with $n$ leaves and a fixed number $k$ of reticulation nodes remains valid in the range $k=o(n^{1/3})$, but not beyond. This is in strong contrast to recently established results for the asymptotic counting formulas for tree-child and normal networks with $n$ leaves and $k$ reticulation nodes, which are valid in the (optimal) range $k=o(n^{1/2})$.
How many cherry-picking sequences are needed to reduce all subtrees of a phylogenetic tree?
Phylogenetic networks are graphs that represent the evolutionary history of species. Recently, the class of orchard phylogenetic networks, which can be reduced by so-called cherry-picking sequences, has gained attention for its computational and biological aspects. In this paper, we study a fundamental question on orchards and their cherry-picking sequences by considering the CoveringNumber problem: given an orchard network $N$, how many cherry-picking sequences are needed to reduce all subnetworks of $N$? We initiate this study by considering the problem for trees. We then show that the covering number can be computed for binary trees recursively using a similar but more fine-grained notion of survival covering number. We also give a recursive formula for the survival covering number of non-binary trees. However, computing the covering number for non-binary trees appears to be considerably more challenging. For this case, we show that the covering number of star trees (whose root is adjacent to all leaves) is equivalent to the so-called SubsetConnectivity problem, which we introduce in this paper. Finally, we show that if there is no restriction on the sequence length, a single sequence of minimum length $\binom{n}{2}$ suffices to reduce all subtrees of a tree on $n$ leaves.
Metropolis-Hastings Sampling of Phylogenetic Networks: Correcting for Symmetries
In phylogenetics, Metropolis-Hastings methods are commonly used to sample phylogenetic trees or networks, for example from Bayesian posteriors. These methods generally use transitions that distinguish all nodes involved, and thus require fully labelled representations of phylogenetic networks. We argue that sampling leaf-labelled phylogenetic networks demands a correction for the number of fully labelled representatives of a leaf-labelled network, or, equivalently, for its internal symmetry. Without correction, there is a danger of undersampling networks with internal symmetries. We show that this correction can be realized by a quotient construction on the Metropolis-Hastings Markov chain, which, in practice, requires the calculation of the size of the network's automorphism group. Using $μ$-vectors, we show that the automorphism group is trivial for orchard networks, and thus also for tree-child networks and trees. This implies that a correction for symmetry is not needed when sampling only from such network classes. More generally, using our Python implementation of the algorithms in this paper, we show that using $μ$-vectors can significantly speed up calculations of automorphism group sizes and thus of Metropolis-Hastings sampling of leaf-labelled networks.
2026-08-04
The maximum quartet distance between phylogenetic trees
The quartet distance counts the four-leaf subsets on which two binary phylogenetic trees display different topologies. We prove that its maximum over trees on $n$ leaves is $(2/3+o(1))\binom{n}{4}$, resolving a conjecture of Bandelt and Dress from 1986 and calibrating the scale of a fundamental metric for comparing phylogenetic trees. The proof reduces arbitrary pairs of trees to caterpillars by means of a common-root planarization and an identity on five-leaf trees.
Enumerating monophyletic characters in mathematical phylogenetics
Grouping species according to their phylogenetic relationships often results in different groups than grouping them according to their shared traits. Monophyletic groups play an important role in this regard, as they are groups of species sharing the same trait and being uniquely defined by a joint phylogenetic subtree. This immediately leads to the question of how to identify possible monophyletic groups in characters, which assign each present-day species a certain trait and which are typically used for phylogenetic tree reconstruction. In our manuscript, we provide a general formula to quantify how many different characters are monophyletic on any given tree and provide simple formulae for binary characters and for certain tree shapes. We also investigate relations between monophyly and the well-known phylogenetic tree reconstruction criterion maximum parsimony by providing a linear-time algorithm which determines the parsimony score together with the monophyly type of a character on a tree.