arXiv++ Combinatorics

Browse math.CO papers from arXiv

species

269 papers tagged with this keyword
2025-10-10 v2
Parameterized Algorithms for Diversity of Networks with Ecological Dependencies
For a phylogenetic tree, the phylogenetic diversity of a set A of taxa is the total weight of edges on paths to A. Finding small sets of maximal diversity is crucial for conservation planning, as it indicates where limited resources can be invested most efficiently. In recent years, efficient algorithms have been developed to find sets of taxa that maximize phylogenetic diversity either in a phylogenetic network or in a phylogenetic tree subject to ecological constraints, such as a food web. However, these aspects have mostly been studied independently. Since both factors are biologically important, it seems natural to consider them together. In this paper, we introduce decision problems where, given a phylogenetic network, a food web, and integers k, and D, the task is to find a set of k taxa with phylogenetic diversity of at least D under the maximize all paths measure, while also satisfying viability conditions within the food web. Here, we consider different definitions of viability, which all demand that a "sufficient" number of prey species survive to support surviving predators. We investigate the parameterized complexity of these problems and present several fixed-parameter tractable (FPT) algorithms. Specifically, we provide a complete complexity dichotomy characterizing which combinations of parameters - out of the size constraint k, the acceptable diversity loss D, the scanwidth of the food web, the maximum in-degree in the network, and the network height h - lead to W[1]-hardness and which admit FPT algorithms. Our primary methodological contribution is a novel algorithmic framework for solving phylogenetic diversity problems in networks where dependencies (such as those from a food web) impose an order, using a color coding approach.
2025-07-30 v2
An explicit power series result for the two type ASEP
Research in combinatorics has often focused on the ASEP (asymmetric simple exclusion process). The ASEP is inspired by processes in statistical mechanics, and involves particles of various species moving around a lattice. The particles do not change species. In the present paper, based on earlier results of Mortimer and Prellberg and others, we obtain a new power series results for the two type ASEP.
Characterizing semi-directed phylogenetic networks and their multi-rootable variants
Published • View Publication • BIB
In evolutionary biology, phylogenetic networks are graphs that provide a flexible framework for representing complex evolutionary histories that involve reticulate evolutionary events. Recently phylogenetic studies have started to focus on a special class of such networks called semi-directed networks. These graphs are defined as mixed graphs that can be obtained by de-orienting some of the arcs in some rooted phylogenetic network, that is, a directed acyclic graph whose leaves correspond to a collection of species and that has a single source or root vertex. However, this definition of semi-directed networks is implicit in nature since it is not clear when a mixed-graph enjoys this property or not. In this paper, we introduce novel, explicit mathematical characterizations of semi-directed networks, and also multi-semi-directed networks, that is, mixed graphs that can be obtained from directed phylogenetic networks that may have more than one root. In addition, through extending foundational tools from the theory of rooted networks into the semi-directed setting - such as cherry picking sequences, omnians, and path partitions - we characterize when a (multi-)semi-directed network can be obtained by de-orienting some rooted network that is contained in one of the well-known classes of tree-child, orchard, tree-based or forest-based networks. These results address structural aspects of (multi-)semi-directed networks and pave the way to improved theoretical and computational analyses of such networks, for example, within the development of algebraic evolutionary models that are based on such networks.
2025-07-12 v2
Counting fixed-point-free Cayley permutations
Published • View Publication • BIB
Two-sort species yield differential equations for functional digraphs of Cayley permutations. From these we obtain an explicit formula for fixed-point-free Cayley permutations and conjecture that their proportion tends to $1/e$, as for permutations and endofunctions. Our approach also yields counting formulas when the functional digraph is a tree, forest, or connected.
2025-06-29 v3
A dichotomy law for certain classes of phylogenetic networks
Published • View Publication • BIB
Many classes of phylogenetic networks have been proposed in the literature. A feature of several of these classes is that if one restricts a network in the class to a subset of its leaves, then the resulting network may no longer lie within this class. This has implications for their biological applicability, since some species -- which are the leaves of an underlying evolutionary network -- may be missing (e.g., they may have become extinct, or there are no data available for them) or we may simply wish to focus attention on a subset of the species. On the other hand, certain classes of networks are `closed' when we restrict to subsets of leaves, such as (i) the classes of all phylogenetic networks or all phylogenetic trees; (ii) the classes of galled networks, simplicial networks, galled trees; and (iii) the classes of networks that have some parameter that is monotone-under-leaf-subsampling (e.g., the number of reticulations, height, etc.) bounded by some fixed value. It is easily shown that a closed subclass of phylogenetic trees is either all trees or a vanishingly small proportion of them (as the number of leaves grows). In this short paper, we explore whether this dichotomy phenomenon holds for other classes of phylogenetic networks, and their subclasses.
2025-06-23
Yet Another Species of Forbidden-distances Chromatic Number
Published in Geombinatorics, vol. 10 no. 3 (2001), pp. 89-95 • Search Publication
This 2001 paper introduces a new type of chromatic number for point sets.
Covariance Decomposition for Distance Based Species Tree Estimation
Published • View Publication • BIB
In phylogenomics, species-tree methods must contend with two major sources of noise; stochastic gene-tree variation under the multispecies coalescent model (MSC) and finite-sequence substitutional noise. Fast agglomerative methods such as GLASS, STEAC, and METAL combine multi-locus information via distance-based clustering. We derive the exact covariance matrix of these pairwise distance estimates under a joint MSC-plus-substitution model and leverage it for reliable confidence estimation, and we algebraically decompose it into components attributable to coalescent variation versus sequence-level stochasticity. Our theory identifies parameter regimes where one source of variance greatly exceeds the other. For both very low and very high mutation rates, substitutional noise dominates, while coalescent variance is the primary contributor at intermediate mutation rates. Moreover, the interval over which coalescent variance dominates becomes narrower as the species-tree height increases. These results imply that in some settings one may legitimately ignore the weaker noise source when designing methods or collecting data. In particular, when gene-tree variance is dominant, adding more loci is most beneficial, while when substitution noise dominates, longer sequences or imputation are needed. Finally, leveraging the derived covariance matrix, we implement a Gaussian-sampling procedure to generate split support values for METAL trees and demonstrate empirically that this approach yields more reliable confidence estimates than traditional bootstrapping.
2025-05-24 v2
Higher Order Bell Symmetric Functions
We study symmetric function analogues of the higher order Bell numbers. Their construction involves iterated plethystic exponential towers mimicking the single variable exponential generating functions for the higher order Bell numbers. We derive explicit recurrence relations for the expansion coefficients of the Bell functions into the monomial and power sum bases of the ring of symmetric functions. Using the machinery of combinatorial species, the Bell functions are proven to be the Frobenius characteristics of the permutation representations of symmetric groups on hyper-partitions of certain orders and sizes. In the order 1 case, we are able to give more details about the expansion coefficients of the Bell functions in terms of vector partitions and divisor sums as well as give a recurrence relation analogous to the well known recursion for the Bell numbers. Lastly, we use Littlewood's reciprocity theorem and the Hardy-Littlewood Tauberian theorem to prove that the Schur expansion coefficients of the order 1 Bell functions are certain asymptotic averages of restriction coefficients.
2025-05-09 v4
Lie-operads and operadic modules from poset cohomology
As observed by Joyal, the cohomology groups of the partition posets are naturally identified with the components of the operad encoding Lie algebras. This connection was explained in terms of operadic Koszul duality by Fresse, and later generalized by Vallette to the setting of decorated partitions. In this article, we set up and study a general formalism which produces a priori operadic structures (operads and operadic modules) on the cohomology of families of posets equipped with some natural recursive structure, that we call "operadic poset species". This framework goes beyond decorated partitions and operadic Koszul duality, and contains the metabelian Lie operad and Kontsevich's operad of trees as two simple instances. In forthcoming work, we will apply our results to the hypertree posets and their connections to post-Lie and pre-Lie algebras.
2025-05-06
Unexpectedly, a symmetry on unlabeled graphs
We exhibit the joint symmetric distribution of the following two parameters on the set of unlabeled, simple, connected graphs with $n$ vertices. The first parameter is the maximal number of leaves attached to a vertex. The second parameter is the size of the largest set of vertices sharing the same closed neighborhood minus $1$. Apparently, this is the first example of a natural, non-trivial equidistribution of graph parameters on unlabeled connected graphs on a fixed set of vertices. Our proof is enumerative, using the theory of species. Exhibiting an explicit bijection interchanging the two parameters remains an open problem.
Coconvex characters on collections of phylogenetic trees
Published • View Publication • BIB
In phylogenetics, a key problem is to construct evolutionary trees from collections of characters where, for a set X of species, a character is simply a function from X onto a set of states. In this context, a key concept is convexity, where a character is convex on a tree with leaf set X if the collection of subtrees spanned by the leaves of the tree that have the same state are pairwise disjoint. Although collections of convex characters on a single tree have been extensively studied over the past few decades, very little is known about coconvex characters, that is, characters that are simultaneously convex on a collection of trees. As a starting point to better understand coconvexity, in this paper we prove a number of extremal results for the following question: What is the minimal number of coconvex characters on a collection of n-leaved trees taken over all collections of size t >= 2, also if we restrict to coconvex characters which map to k states? As an application of coconvexity, we introduce a new one-parameter family of tree metrics, which range between the coarse Robinson-Foulds distance and the much finer quartet distance. We show that bounds on the quantities in the above question translate into bounds for the diameter of the tree space for the new distances. Our results open up several new interesting directions and questions which have potential applications to, for example, tree spaces and phylogenomics.
2025-03-06
A New Representation of Ewens-Pitman's Partition Structure and Its Characterization via Riordan Array Sums
Ewens-Pitman's partition structure arises as a system of sampling consistent probability distributions on set partitions induced by the Pitman-Yor process. It is widely used in statistical applications, particularly in species sampling models in Bayesian nonparametrics. Drawing references from the area of representation theory of the infinite symmetric group, we view Ewens-Pitman's partition structure as an example of a non-extreme harmonic function on a branching graph, specifically, the Kingman graph. Taking this perspective enables us to obtain combinatorial and algebraic constructions of this distribution using the interpolation polynomial approach proposed by Borodin and Olshanski (The Electronic Journal of Combinatorics, 7, 2000). We provide a new explicit representation of Ewens-Pitman's partition structure using modern umbral interpolation based on Sheffer polynomial sequences. In addition, we show that a certain type of marginals of this distribution can be computed using weighted row sums of a Riordan array. In this way, we show that some summary statistics and estimators derived from Ewens-Pitman's partition structure can be obtained using methods of generating functions. This approach simplifies otherwise cumbersome calculations of these quantities often involving various special combinatorial functions. In addition, it has the added benefit of being amenable to symbolic computation.
2025-03-02 v3
Multispecies inhomogeneous $t$-PushTASEP from antisymmetric fusion
Published in Electron. J. Probab. 30: 1-28 (2025) • View Publication • BIB
We investigate the recently introduced inhomogeneous $n$-species $t$-PushTASEP, a long-range stochastic process on a periodic lattice. A Baxter-type formula is established, expressing the Markov matrix as an alternating sum of commuting transfer matrices over all the fundamental representations of $U_t(\widehat{sl}_{n+1})$. This superposition acts as an inclusion-exclusion principle, selectively extracting the sequential particle transitions characteristic of the PushTASEP, while canceling forbidden channels. The homogeneous specialization connects the PushTASEP to ASEP, showing that the two models share eigenstates and a common integrability structure.
Orthology and Near-Cographs in the Context of Phylogenetic Networks
Published • View Publication • BIB
Orthologous genes, which arise through speciation, play a key role in comparative genomics and functional inference. In particular, graph-based methods allow for the inference of orthology estimates without prior knowledge of the underlying gene or species trees. This results in orthology graphs, where each vertex represents a gene, and an edge exists between two vertices if the corresponding genes are estimated to be orthologs. Orthology graphs inferred under a tree-like evolutionary model must be cographs. However, real-world data often deviate from this property, either due to noise in the data, errors in inference methods or, simply, because evolution follows a network-like rather than a tree-like process. The latter, in particular, raises the question of whether and how orthology graphs can be derived from or, equivalently, are explained by phylogenetic networks. Here, we study the constraints imposed on orthology graphs when the underlying evolutionary history follows a phylogenetic network instead of a tree. We show that any orthology graph can be represented by a sufficiently complex level-k network. However, such networks lack biologically meaningful constraints. In contrast, level-1 networks provide a simpler explanation, and we establish characterizations for level-1 explainable orthology graphs, i.e., those derived from level-1 evolutionary histories. To this end, we employ modular decomposition, a classical technique for studying graph structures. Specifically, an arbitrary graph is level-1 explainable if and only if each primitive subgraph is a near-cograph (a graph in which the removal of a single vertex results in a cograph). Additionally, we present a linear-time algorithm to recognize level-1 explainable orthology graphs and to construct a level-1 network that explains them, if such a network exists.
2025-01-16 v3
Predicting the depth of the most recent common ancestor of a random sample of $k$ species: the impact of phylogenetic tree shape
Published • View Publication • BIB
We consider the following question: how close to the ancestral root of a phylogenetic tree is the most recent common ancestor of $k$ species randomly sampled from the tips of the tree? For trees having shapes predicted by the Yule-Harding model, it is known that the most recent common ancestor is likely to be close to (or equal to) the root of the full tree, even as $n$ becomes large (for $k$ fixed). However, this result does not extend to models of tree shape that more closely describe phylogenies encountered in evolutionary biology. We investigate the impact of tree shape (via the Aldous $β-$splitting model) to predict the number of edges that separate the most recent common ancestor of a random sample of $k$ tip species and the root of the parent tree they are sampled from. Both exact and asymptotic results are presented. We also briefly consider a variation of the process in which a random number of tip species are sampled.
Species of Rota-Baxter algebras by rooted trees, twisted bialgebras and Fock functors
As a fundamental and ubiquitous combinatorial notion, species has attracted sustained interest, generalizing from set-theoretical combinatorial to algebraic combinatorial and beyond. The Rota-Baxter algebra is one of the algebraic structures with broad applications from Renormalization of quantum field theory to integrable systems and multiple zeta values. Its interpretation in terms of monoidal categories has also recently appeared. This paper studies species of Rota-Baxter algebras, making use of the combinatorial construction of free Rota-Baxter algebras in terms of angularly decorated trees and forests. The notion of simple angularly decorated forests is introduced for this purpose and the resulting Rota-Baxter species is shown to be free. Furthermore, a twisted bialgebra structure, as the bialgebra for species, is established on this free Rota-Baxter species. Finally, through the Fock functor, another proof of the bialgebra structure on free Rota-Baxter algebras is obtained.
Integro-differential rings on species and derived structures
In the theory of species, differential as well as integral operators are known to arise in a natural way. In this paper, we shall prove that they precisely fit together in the algebraic framework of integro-differential rings, which are themselves an abstraction of classical calculus (incorporating its Fundamental Theorem). The results comprise (set) species as well as linear species. Localization of (set) species leads to the more general structure of modified integro-differential rings, previously employed in the algebraic treatment of Volterra integral equations. Furthermore, the ring homomorphism from species to power series via taking generating series is shown to be a (modified) integro-differential ring homomorphism. As an application, a topology and further algebraic operations are imported to virtual species from the general theory of integro-differential rings.
2024-11-13 v2
Enumerative aspects of Caylerian polynomials
Eulerian polynomials record the distribution of descents over permutations. Caylerian polynomials likewise record the distribution of descents over Cayley permutations, where a Cayley permutation is a word of positive integers such that if a number appears in the word then all positive integers less than that number also appear in the word. Using combinatorial species and sign-reversing involutions we derive counting formulas and generating functions for the Caylerian polynomials as well as for related refined polynomials.
2024-11-01 v3
Simplifying and Characterizing DAGs and Phylogenetic Networks via Least Common Ancestor Constraints
Published • View Publication • BIB
Rooted phylogenetic networks, or more generally, directed acyclic graphs (DAGs), are widely used to model species or gene relationships that traditional rooted trees cannot fully capture, especially in the presence of reticulate processes or horizontal gene transfers. Such networks or DAGs are typically inferred from observable data (e.g. genomic sequences of extant species), providing only an estimate of the true evolutionary history. However, these inferred DAGs are often complex and difficult to interpret. In particular, many contain vertices that do not serve as least common ancestors (LCAs) for any subset of the underlying genes or species, thus may lack direct support from the observable data. In contrast, LCA vertices are witnessed by historical traces justifying their existence and thus represent ancestral states substantiated by the data. To reduce unnecessary complexity and eliminate unsupported vertices, we aim to simplify a DAG to retain only LCA vertices while preserving essential evolutionary information. In this paper, we characterize $\mathrm{LCA}$-relevant and $\mathrm{lca}$-relevant DAGs, defined as those in which every vertex serves as an LCA (or unique LCA) for some subset of taxa. We introduce methods to identify LCAs in DAGs and efficiently transform any DAG into an $\mathrm{LCA}$-relevant or $\mathrm{lca}$-relevant one while preserving key structural properties of the original DAG or network. This transformation is achieved using a simple operator ``$\ominus$'' that mimics vertex suppression.
2024-08-30
Characterising rooted and unrooted tree-child networks
Rooted phylogenetic networks are used by biologists to infer and represent complex evolutionary relationships between species that cannot be accurately explained by a phylogenetic tree. Tree-child networks are a particular class of rooted phylogenetic networks that has been extensively investigated in recent years. In this paper, we give a novel characterisation of a tree-child network $\mathcal{R}$ in terms of cherry-picking sequences that are sequences on the leaves of $\mathcal{R}$ and reduce it to a single vertex by repeatedly applying one of two reductions to its leaves. We show that our characterisation extends to unrooted tree-child networks which are mostly unexplored in the literature and, in turn, also offers a new approach to settling the computational complexity of deciding if an unrooted phylogenetic network can be oriented as a rooted tree-child network.