arXiv++ Combinatorics

Browse math.CO papers from arXiv

q-bio.PE ↗ arXiv

21 papers in this category
Encoding level-3 semi-directed phylogenetic networks by quarnets and quinnets
Phylogenetic networks generalize phylogenetic trees as models of evolutionary history, allowing lineages to merge as well as to diverge. For many types of genetic data the root position of such a network cannot be recovered, so that only a semi-directed network can be inferred: a mixed graph in which only the edges entering a reticulation vertex are directed. A common strategy for inferring such a network is to first infer the subnetwork it induces on each set of $k\geq 3$ of its leaves, called a $k$-net, and then to assemble these pieces. This can only succeed if the $k$-nets determine the network, in which case that network is said to be encoded by its $k$-nets. Semi-directed networks of level-1 and 2, those whose biconnected components contain at most one, respectively two, reticulations, are known to be encoded by their $4$-nets, or quarnets, whereas level-3 networks are not. Even so, in this paper we show that level-3 semi-directed networks are encoded by their $5$-nets, or quinnets, and we characterize the limitation of quarnets exactly: we show that a single previously reported counterexample captures the only obstruction, every other level-3 network being encoded by its quarnets. Our proofs rest on a collection of encoding results for individual structural features of a network, which we establish for networks of arbitrary level and which are of independent interest.
2026-10-02
Exact and asymptotic enumeration of unrestricted binary phylogenetic networks through automorphism weights
Let $\cP_{\ell,k}$ be the set of rooted binary phylogenetic networks with $\ell$ labelled leaves, $k$ reticulations and no parallel edges. We write $|\cP_{\ell,k}|=W_k(\ell)+D_k(\ell)$, where the weighted count $W_k(\ell)$ adds the inverse orders of the leaf-fixing automorphism groups and the defect $D_k(\ell)$ collects the remainder. The weighted count satisfies, for every $k$, a recursion over the source layers of the tree-component structure that involves neither a list of component graphs nor any distinction between symmetric and asymmetric configurations, and its exponential generating function is a Laurent polynomial in $\sqrt{1-2x}$. Automorphism groups of networks are $2$-groups, elementary abelian for $k\le5$ but not in general. For $k\le5$ the defect is the weighted count of networks with a distinguished involution, which obeys an extension of the same recursion. An exact symbolic evaluation of the two recursions yields $|\cP_{\ell,k}|$ in closed form for $k\le5$, the case $k=5$ being new; it reproduces the published counts for $k\le4$ and corrects a coefficient in a published generating function for $k=3$. For every $k$, uniformly over explicit ranges of $k$, we prove that the non-tree-child networks are a fraction $2k(k-1)/\ell$ of the tree-child networks to leading order, which gives the third term of the asymptotic expansion of $|\cP_{\ell,k}|$. We also prove that reticulation-visible networks exceed tree-child networks by the fraction $k(k-1)/\ell$, and that a uniformly random network in $\cP_{\ell,k}$ has a nontrivial automorphism with probability $k(k-1)/(4\ell^3)$ to leading order.
2026-10-02 v3
Exact Enumeration of Phylogenetic Networks: The Tree-Child, Reticulation-Visible and Orchard Hierarchy
We develop a unified framework for the exact enumeration and asymptotic analysis of the three most studied classes of phylogenetic networks: tree-child (TC), reticulation-visible (RV) and orchard networks, whose cardinalities satisfy the strict ordering $|\mathrm{TC}_{\ell,k}|<|\mathrm{RV}_{\ell,k}|<|\mathrm{Orch}_{\ell,k}|$ for reticulation number $k\geq2$ (with $\mathrm{TC}\subsetneq\mathrm{RV}$ and $\mathrm{TC}\subsetneq\mathrm{Orch}$, while $\mathrm{RV}$ and $\mathrm{Orch}$ are incomparable as sets). Using the Chang--Fuchs structural theorem, we derive a two-level master functional equation for the RV bivariate generating function and obtain exact closed-form identities for the differences $Δ_k(\ell):=|RV_{\ell,k}|-|TC_{\ell,k}|$ for $k=2,3$, with the asymptotic universality $Δ_k(\ell)/|TC_{\ell,k}|\sim k!/\ell$. For orchard networks, we prove a \emph{universal hypergeometric law} that resolves the exact enumeration problem for all $\ell$: the column generating function $F_\ell(v)$ is rational with denominator $D_\ell(v)=\prod_{j=2}^\ell X_j(v)$, where \[ X_\ell(v) = \sum_{k=0}^{\lfloor\ell/2\rfloor}(-1)^k\, \frac{\ell!}{(\ell-2k)!\,k!}\,v^k \] is the matching polynomial of the complete graph $K_\ell$ and a rescaled Jacobi polynomial. This immediately resolves the intractable $\ell=9$ case: $D_9$ has degree 20, dominant growth rate $\approx40.73$, and all spectral roots are positive real. A complete enumeration table is provided extending the published data of Cardona, Ribas and Pons.
2026-10-01 v2
Characterisations of Planar Galled Networks
Rooted phylogenetic networks are widely used to represent the evolution of species that have undergone reticulate processes. However, these networks can be highly non-planar, making them more difficult to visualise and interpret than evolutionary trees. In this paper, we investigate planarity properties of galled networks, an important subclass of phylogenetic networks. We show that all planar galled networks are necessarily upward planar. Furthermore, by leveraging recent results on planar phylogenetic networks, we provide three characterisations for each of the outerplanar and terminal planar galled network classes in terms of forbidden vertex configurations, forbidden directed subgraphs, and forbidden structures in their associated underlying undirected graphs. These results contribute to a deeper understanding of the structural properties of galled networks and may inform future methods for their construction and visualisation.
2026-10-01 v2
Exact Counts of Binary Phylogenetic Networks with Four Reticulations
Phylogenetic networks provide a flexible framework for representing reticulate evolutionary processes, such as hybridization, introgression, recombination, and horizontal gene transfer. However, their combinatorial complexity makes even basic enumeration problems difficult. Building on our previous work for networks with up to three reticulations, we derive an explicit closed-form formula for the number of unrestricted rooted binary phylogenetic networks with four reticulations on \(n\) labeled taxa. Our approach is based on tree-component graphs. We classify the 79 possible component graphs corresponding to networks with four reticulations into ten groups. We then enumerate the networks associated with each group by combining known counts of one-component networks, forests, and networks with fewer reticulations. Summing these contributions yields the desired formula. This result extends the exact enumeration of unrestricted binary phylogenetic networks to four reticulations and further demonstrates the effectiveness of component graphs for systematically organizing and counting increasingly complex network classes.
Characterization of tree-child networks in terms of mu vectors
We characterize tree-child phylogenetic networks in terms of their mu-representations. First, we give a structural characterization of tree-child networks by means of ordered tree-path decompositions. We then translate this decomposition into a set of purely vectorial conditions on finite subsets M in N^n. We prove that such a set M is the mu-representation of a tree-child phylogenetic network if and only if it is tree-child mu-compatible. This provides a feasibility criterion for tree-child mu-representations which can be used as a basis for reconstruction and further algorithmic applications. Note that this paper presents results arising from ongoing research on tree-child networks and that the results will be further developed and placed into proper context in subsequent versions.
2026-09-11
NOC NOC, who's there? Clustering systems of tree-child and normal networks
Clustering systems provide a natural way to encode structural information contained in phylogenetic networks. In this note, we study the clustering systems of normal and tree-child networks through an overlap-based property of set systems, called not-overlap-covered (NOC). We show that the NOC property is equivalent to inclusion-visibility, a memberwise formulation of the strict-compatibility condition previously used for tree-child clustering systems. We characterize normal networks as precisely the semi-regular networks whose clustering systems satisfy NOC. Consequently, a clustering system is realized by a normal network if and only if it satisfies NOC, or equivalently, if every one of its clusters is inclusion-visible. In this case, the Hasse diagram provides a canonical normal realization. These are exactly the clustering systems realized by tree-child networks. The NOC formulation yields a sharp quadratic upper bound on the number of distinct clusters of tree-child and normal networks and a direct polynomial-time recognition algorithm. Finally, we explore several consequences of the NOC perspective beyond the phylogenetic setting. These include connections to the enumeration of normal networks, an order-theoretic interpretation of inclusion-visibility, structural properties of NOC set systems, and a tractable special case of Minimum Set Cover, which is NP-hard in general.
2026-09-09 v2
Regularizing and Normalizing DAGs and Phylogenetic Networks
Phylogenetic networks and, more generally, directed acyclic graphs (DAGs) represent hierarchical structure beyond trees, for instance in the presence of reticulate evolutionary events such as hybridization or horizontal gene transfer. A central question is which parts of such graphs are essential with respect to leaf-observable information, and which parts can be removed without changing this information. Resolving this question can lead to principled simplification methods for phylogenetic networks, such as the recent normalization approach of Francis et al. In this paper, we study this question from three related perspectives: clusters displayed by a DAG $G$, least common ancestors (LCAs) of subsets of its leaf set, and visibility, a path-based property of vertices. We first introduce an LCA-based simplification procedure called $i$-regularization. For a DAG $G$ and $i\geq 1$, the DAG $\reg_i(G)$ retains precisely those vertices that occur as unique LCAs of leaf subsets of size at most $i$, removes the remaining non-leaf vertices by a graph-editing operation $\ominus$, and then deletes shortcuts. We show that $\reg_i(G)$ admits a Hasse-diagram characterization in terms of the corresponding lca-clusters. We then compare LCA-based regularization with normalization. Using the same $\ominus$-operator, we describe the cover construction underlying normalization, identify visible vertices that are nevertheless removed, and characterize when regularization and normalization coincide. Together, these results provide a unified framework for cluster-based, LCA-based, and visibility-based simplifications of DAGs and phylogenetic networks.
2026-09-04
A Short Combinatorial Proof of the Pons-Batle Identity for Counting Tree-Child Networks
Tree-child networks are a useful class of binary phylogenetic networks. The Pons--Batle identity (Pons and Batle, \textit{Scientific Reports}, 2021) states that the number $a_{n,k}$ of tree-child networks with $k$ reticulations on $n$ taxa satisfies \[ a_{n,k}=(n-k+1)a_{n,k-1} +\frac{n(2n+k-3)}{n-k}a_{n-1,k}. \] In this paper, we present a short combinatorial proof of this identity.
2026-08-26
Tree Buckets and the Reconstruction of Pairs of Phylogenetic Trees
Phylogenetic trees are used in evolutionary biology to represent the evolutionary history of a collection of taxa. As we have incomplete information about any evolutionary history, recovering trees from partial information is a focus of phylogenetic combinatorics. However, in some cases the available data does not describe a single phylogenetic tree. We consider recovery of pairs of phylogenetic trees from their combined subtrees with $k$ leaves, which we call a $k$-bucket. We establish the exact cases in which these pairs of trees are recoverable from their subtrees with a single leaf removed, both when just considering the structure of the trees, and when additionally considering the set of taxa on the leaves. We also consider recovery of pairs of trees with labelled leaves from their rooted triples, and establish that they are recoverable up to a sequence of subtree swaps.
2026-08-24
On the maximum size of 2-weakly compatible split systems
We consider a Turán-type problem arising in phylogenetics: determining the maximum size of a 2-weakly compatible split system. This compatibility condition arises in the reconstruction of phylogenetic networks from quartet weights. It was previously shown that a 2-weakly compatible split system has size at most \[ 3\binom{n}{4}+\binom{n}{2}. \] We prove that the maximum size is $O(n^{5/2})$.
Enumerating monophyletic characters in mathematical phylogenetics
Grouping species according to their phylogenetic relationships often results in different groups than grouping them according to their shared traits. Monophyletic groups play an important role in this regard, as they are groups of species sharing the same trait and being uniquely defined by a joint phylogenetic subtree. This immediately leads to the question of how to identify possible monophyletic groups in characters, which assign each present-day species a certain trait and which are typically used for phylogenetic tree reconstruction. In our manuscript, we provide a general formula to quantify how many different characters are monophyletic on any given tree and provide simple formulae for binary characters and for certain tree shapes. We also investigate relations between monophyly and the well-known phylogenetic tree reconstruction criterion maximum parsimony by providing a linear-time algorithm which determines the parsimony score together with the monophyly type of a character on a tree.
Classes of phylogenetic networks that are robust to root placement
Standard phylogenetic reconstruction techniques often yield unrooted phylogenetic networks; these are subsequently rooted to infer evolutionary history. A common problem in this process is to determine the structural classes to which the resulting network will belong. In this paper, we investigate unrooted networks in which the choice of any root results in a valid rooted phylogenetic network, a property we define as {\em robustly orientable}. We then establish a strict structural condition for this class, specifically, that an unrooted network is robustly orientable if and only if it contains no sink components. We also show that if an unrooted network is level-$2$ or less, or if it is tree-based, then it is robustly orientable. Furthermore, we define an unrooted network to be {\em robustly class $\mathcal C$} if the choice of any root results in a network belonging to class $\mathcal C$. We demonstrate that an unrooted network is robustly tree-child or robustly stack-free if and only if it is level-$1$ or less. Finally, we show that a phylogenetic network is robustly normal if and only if it is a phylogenetic tree.
Proximity Measures for Classes of Phylogenetic Networks
Phylogenetic networks are used to represent the evolutionary history of species. Due to biological interpretations and computational advantages, researchers have focused on restricted classes of phylogenetic networks, such as tree-child, orchard, and tree-based. These classes capture different notions of tree-likeness: tree-child networks require every internal vertex to have a taxon reachable by a tree path, orchard networks are trees with horizontal arcs (for modelling histories rife with horizontal gene transfers), and tree-based networks are trees with additional (not-necessarily horizontal) arcs. A natural question to ask is ``how far is a given network from belonging to a particular class?'' This motivates the study of proximity measures, which measure the minimum number of graph modifications required to transform a network into one belonging to a particular class. In this paper, we consider three proximity measures based on leaf addition, valid arc deletion, and arc deletion. We study pairwise comparability of the proximity measures, prove complexity results, and derive extremal bounds for the classes of tree, tree-child, orchard, and tree-based networks.
2026-07-06
Polynomial encoding of rooted trees with branch lengths
Phylogenetic trees are rooted trees with branch lengths that record genetic divergence or elapsed time, and quantifying differences between them is central to a wide range of evolutionary and epidemiological analyses. Graph-polynomial encodings of rooted trees provide an accurate, interpretable, and computationally efficient way to compare tree shapes, but existing polynomial encodings must be paired with auxiliary structures to study rooted trees with branch lengths. We introduce a bivariate polynomial encoding that incorporates branch lengths directly into a recursive computation from the leaf vertices to the root vertex of a tree. We prove that, for rooted trees with branch lengths and no vertices of degree two, which include all standard phylogenetic trees, two trees have the same polynomial if and only if their underlying unlabeled trees are isomorphic and the branch lengths of corresponding edges are equal. We apply the polynomial encoding to three published HIV-1 phylogenies sampled in different epidemiological settings and show that it accurately separates the three datasets based on their tree topologies and branch lengths, outperforming previous polynomial-based approaches for analyzing rooted trees with branch lengths.
A parameterized family of balance indices for phylogenetic networks
We introduce a new family of balance indices for phylogenetic networks: the $H_α$ indices, where $α$ is a positive real number. This family includes the $B_2$ index as a special case ($α= 1$) and provides a natural extension of the Sackin index to phylogenetic networks. We show that the $H_α$ indices share many structural properties with the $B_2$ index, most notably a "grafting property" that makes it possible to express the $H_α$ index of a network in terms of the $H_α$ indices of its biconnected components. These properties allow us to identify networks that minimize / maximize $H_α$ for various classes of phylogenetic networks, and to study its distribution for several models of random trees and networks (in particular, Galton-Watson trees and binary Markov branching trees, with a focus on the Yule and PDA models). Finally, we show how local limits can be used to analyze the asymptotic behavior of $H_α$ for large trees and networks, and we obtain general results for the moments of $H_α$ for a broad class of random phylogenetic networks known as blowups of Galton-Watson trees.
2026-06-12
Note on the Maximum Number of Trees Displayed by a Tree-Child Network
In this note, we show that, for all $n\ge 2$, the number of distinct rooted binary phylogenetic $X$-trees displayed by a binary tree-child network $\mathcal{N}$ on $X$ with $n$ leaves is at most $2^{n-1}-1$ and that this upper bound is sharp. Furthermore, if $\mathcal{N}$ displays exactly $2^{n-1}-1$ such trees, then exactly one rooted binary phylogenetic $X$-tree is displayed twice, and this tree can be canonically found by iteratively replacing a reticulated cherry with a cherry.
2026-05-07
A $μ$-distance for semidirected orchard phylogenetic networks
In evolutionary biology, phylogenetic networks are now widely used to represent the historical relationships between species and population, when this history includes reticulation events such as hybridization, gene flow and admixture between populations. Semidirected phylogenetic networks are appropriate models when the direction of some edges and the root position are not identifiable from data. Comparing semidirected networks is important in many applications. For rooted and directed networks, a $μ$-representation was originally introduced to distinguish tree-child networks, and has since been extended in two different directions: to the larger class of orchard directed networks by adding an extra component that counts paths to reticulations; and to semidirected networks, through an edge-based variant. However, the latter does not provide a distance between semidirected and orchard networks. We introduce here a new edge-based $μ$-representation capable of distinguishing distinct orchard binary semidirected networks. For this class, we provide a reconstruction algorithm and therefore obtain a true distance that is computable in polynomial time.
2026-04-06
Nested tree space: a geometric framework for co-phylogeny
Nested (or reconciled) phylogenetic trees model co-evolutionary systems in which one evolutionary history is embedded within another. We introduce a geometric framework for such systems by defining $σ$-space, a moduli space of fully nested ultrametric phylogenetic trees with a fixed leaf map. Generalizing the $τ$-space of Gavryushkin and Drummond, $σ$-space is constructed as a cubical complex parametrised by nested ranked tree topologies and inter-event time coordinates of the combined host and parasite speciation events. We characterise admissible orderings via binary \textit{nesting sequences} and organise them into a natural poset. We show that $σ$-space is contractible and satisfies Gromov's cube condition, and is therefore CAT(0). In particular, it admits unique geodesics and well-defined Fréchet means. We further describe its geometric structure, including boundary strata corresponding to cospeciation events, and relate it to products of ultrametric tree spaces via natural forgetful maps.
A Class of Unrooted Phylogenetic Networks Inspired by the Properties of Rooted Tree-Child Networks
A directed phylogenetic network is tree-child if every non-leaf vertex has a child that is not a reticulation. As a class of directed phylogenetic networks, tree-child networks are very useful from a computational perspective. For example, several computationally difficult problems in phylogenetics become tractable when restricted to tree-child networks. At the same time, the class itself is rich enough to contain quite complex networks. Furthermore, checking whether a directed network is tree-child can be done in polynomial time. In this paper, we seek a class of undirected phylogenetic networks that is rich and computationally useful in a similar way to the class tree-child directed networks. A natural class to consider for this role is the class of tree-child-orientable networks which contains all those undirected phylogenetic networks whose edges can be oriented to create a tree-child network. However, we show here that recognizing such networks is NP-hard, even for binary networks, and as such this class is inappropriate for this role. Towards finding a class of undirected networks that fills a similar role to directed tree-child networks, we propose new classes called $q$-cuttable networks, for any integer $q\geq 1$. We show that these classes have many of the desirable properties, similar to tree-child networks in the rooted case, including being recognizable in polynomial time, for all $q\geq 1$. Towards showing the computational usefulness of the class, we show that the NP-hard problem Tree Containment is polynomial-time solvable when restricted to $q$-cuttable networks with $q\geq 3$.