arXiv++ Combinatorics

Browse math.CO papers from arXiv

species

269 papers tagged with this keyword
Relative Timing Information and Orthology in Evolutionary Scenarios
Published • View Publication • BIB
Evolutionary scenarios describing the evolution of a family of genes within a collection of species comprise the mapping of the vertices of a gene tree $T$ to vertices and edges of a species tree $S$. The relative timing of the last common ancestors of two extant genes (leaves of $T$) and the last common ancestors of the two species (leaves of $S$) in which they reside is indicative of horizontal gene transfers (HGT) and ancient duplications. Orthologous gene pairs, on the other hand, require that their last common ancestors coincides with a corresponding speciation event. The relative timing information of gene and species divergences is captured by three colored graphs that have the extant genes as vertices and the species in which the genes are found as vertex colors: the equal-divergence-time (EDT) graph, the later-divergence-time (LDT) graph and the prior-divergence-time (PDT) graph, which together form an edge partition of the complete graph. Here we give a complete characterization in terms of informative and forbidden triples that can be read off the three graphs and provide a polynomial time algorithm for constructing an evolutionary scenario that explains the graphs, provided such a scenario exists. We show that every EDT graph is perfect. While the information about LDT and PDT graphs is necessary to recognize EDT graphs in polynomial-time for general scenarios, this extra information can be dropped in the HGT-free case. However, recognition of EDT graphs without knowledge of putative LDT and PDT graphs is NP-complete for general scenarios. In contrast, PDT graphs can be recognized in polynomial-time. We finally connect the EDT graph to the alternative definitions of orthology that have been proposed for scenarios with horizontal gene transfer. With one exception, the corresponding graphs are shown to be colored cographs.
2022-11-14 v2
Directed hereditary species and decomposition spaces
We introduce the notion of directed hereditary species and show that they have associated monoidal decomposition spaces, comodule bialgebras, and operadic categories. The notion subsumes Schmitt's hereditary species, Gálvez--Kock--Tonks directed restrictions species, and a directed version of Carlier's construction of monoidal decomposition spaces and comodule bialgebras. In addition to all the examples of Schmitt, Gálvez--Kock--Tonks and Carlier, the new construction covers also the Fauvet--Foissy--Manchon comodule bialgebra of finite topological spaces, the Calaque--Ebrahimi-Fard--Manchon comodule bialgebra of rooted trees, and the Faà di Bruno comodule bialgebra of linear trees.
2022-11-11 v2
A $\mathrm{GL}(\mathbb{F}_q)$-compatible Hopf algebra of unitriangular class functions
Published • View Publication • BIB
This paper constructs a novel Hopf algebra $\mathsf{cf}(\mathrm{UT}_{\bullet})$ on the class functions of the unipotent upper triangular groups $\mathrm{UT}_{n}(\mathbb{F}_{q})$ over a finite field. This construction is representation theoretic in nature and uses the machinery of Hopf monoids in the category of vector species. In contrast with a similar known construction, this Hopf algebra has the property that induction to the finite general linear group induces a homomorphism to Zelevinsky's Hopf algebra of $\mathrm{GL}_{n}(\mathbb{F}_{q})$ class functions. Furthermore, $\mathsf{cf}(\mathrm{UT}_{\bullet})$ contains a Hopf subalgebra which is isomorphic to a known combiantorial Hopf algebra, previously used to prove a conjecture about chromatic quasisymmetric functions. Some additional Hopf algebraic properties are also established.
2022-10-27
Antipode formulas for pattern Hopf algebras
The permutation pattern Hopf algebra is a commutative filtered and connected Hopf algebra. Its product structure stems from counting patterns of a permutation, interpreting the coefficients as permutation quasi-shuffles. The Hopf algebra was shown to be a free commutative algebra and to fit into a general framework of pattern Hopf algebras, via species with restrictions. In this paper we introduce the cancellation-free and grouping-free formula for the antipode of the permutation pattern Hopf algebra. To obtain this formula, we use the popular sign-reversing involution method, by Benedetti and Sagan. This formula has applications on polynomial invariants on permutations, in particular for obtaining reciprocity theorems. On our way, we also introduce the packed word patterns Hopf algebra and present a formula for its antipode. Other pattern algebras are discussed here, notably on parking functions, which recovers notions recently studied by Adeniran and Pudwell, and by Qiu and Remmel.
2022-06-09
Free pre-lie algebras of finite posets
Published • View Publication • BIB
In this paper, we first recall the construction of a twisted pre-Lie algebra structure on the species of finite connected topological spaces. Then we construct the corresponding nonassociative permutative coproduct, and we prove that the vector space generated by isomorphism classes of finite posets is a free pre-Lie algebra and is a co-free non-associative permutative coalgebra. In the end, we give an explicit duality between the non-associative permutative product and the proposed non-associative permutative coproduct. Finally, we prove that the results in this paper remain true for the finite connected topological spaces.
A Model for Birdwatching and other Chronological Sampling Activities
Published • View Publication • BIB
In many real life situations one has $m$ types of random events happening in chronological order within a time interval and one wishes to predict various milestones about these events or their subsets. An example is birdwatching. Suppose we can observe up to $m$ different types of birds during a season. At any moment a bird of type $i$ is observed with some probability. There are many natural questions a birdwatcher may have: how many observations should one expect to perform before recording all types of birds? Is there a time interval where the researcher is most likely to observe all species? Or, what is the likelihood that several species of birds will be observed at overlapping time intervals? Our paper answers these questions using a new model based on random interval graphs. This model is a natural follow up to the famous coupon collector's problem.
2022-05-11 v4
Hopf monoids of set families
A \textit{grounded set family} on $I$ is a subset $F\subseteq2^I$ such that $\emptyset\in F$. We study a linearized Hopf monoid \textbf{SF} on grounded set families, with restriction and contraction inspired by the corresponding operations for antimatroids. Many known combinatorial species, including simplicial complexes and matroids, form Hopf submonoids of \textbf{SF}, although not always with the "standard" Hopf structure (for example, our contraction operation is not the usual contraction of matroids). We use the topological methods of Aguiar and Ardila to obtain a cancellation-free antipode formula for the Hopf submonoid of lattices of order ideals of finite posets. Furthermore, we prove that the Hopf algebra of lattices of order ideals of chain gangs extends the Hopf algebra of symmetric functions, and that its character group extends the group of formal power series in one variable with constant term 1 under multiplication.
2022-03-18 v2
Planar Rooted Phylogenetic Networks
Published • View Publication • BIB
A rooted phylogenetic network is a directed acyclic graph with a single root, whose sinks correspond to a set of species. As such networks are useful for representing the evolution of species that have undergone reticulate evolution, there has been great interest in developing the theory behind and algorithms for constructing them. However, unlike evolutionary trees, these networks can be highly non-planar, which can make them difficult to visualise and interpret. Here we investigate properties of planar rooted phylogenetic networks and algorithms for deciding whether or not rooted networks have certain special planarity properties. In particular, we introduce three natural subclasses of planar rooted phylogenetic networks and show that they form a hierarchy. In addition, for the well-known level-k networks, we show that level-1, -2, -3 networks are always outer, terminal, and upward planar, respectively, and that level-4 networks are not necessarily planar. Finally, we show that a regular network is terminal planar if and only if it is pyramidal. Our results make use of the highly developed field of planar digraphs, and we believe that the link between phylogenetic networks and planar graphs should prove useful in future for developing new approaches to both construct and visualise phylogenetic networks.
2022-03-13 v2
Joint $q$-moments and shift invariance for the multi-species $q$-TAZRP on the infinite line
This paper presents a novel method for computing certain particle locations in the multi-species $q$-TAZRP (totally asymmetric zero range process). The method is based on a decomposition of the process into its discrete-time embedded Markov chain, which is described more generally as a monotone process on a graded partially ordered set; and an independent family of exponential random variables. A further ingredient is explicit contour integral formulas for the transition probabilities of the $q$-TAZRP. The main result of this method is a shift invariance for the multi-species $q$-TAZRP on the infinite line. By a previously known Markov duality result, these particle locations are the same as joint $q$-moments. One particular special case is that for step initial conditions, ordered multi-point joint $q$-moments of the $n$-species $q$-TAZRP match the $n$-point joint $q$-moments of the single-species $q$-TAZRP. Thus, we conjecture that the Airy$_2$ process describes the joint multi-point fluctuations of multi-species $q$-TAZRP. As a probabilistic application of this result, we find explicit contour integral formulas for the joint $q$-moments of the multi-species $q$-TAZRP in the diffusive scaling regime.
2022-02-20
A Linear Time, and Constant Space, Algorithm to Compute the Mixed Moments of the Multivariate Normal Distributions
Using recurrences gotten from the Apagodu-Zeilberger Multivariate Almkvist-Zeilberger algorithm we present a linear-time, and constant-space, algorithm to compute the general mixed moments of the k-variate general normal distribution, with any covariance matrix, for any specific k. Besides their obvious importance in statistics, these numbers are also very significant in enumerative combinatorics, since they count in how many ways, in a species with k different genders, a bunch of individuals can all get married, keeping track of the different kinds of heterosexual marriages. We completely implement our algorithm (with an accompanying Maple package, MVNM.txt) for the bivariate and trivariate cases (and hence taking care of our own 2-sex society and a putative 3-sex society), but alas, the actual recurrences for larger k took too long for us to compute. We leave them as computational challenges.
Forest-based networks
Published • View Publication • BIB
In evolutionary studies it is common to use phylogenetic trees to represent the evolutionary history of a set of species. However, in case the transfer of genes or other genetic information between the species or their ancestors has occurred in the past, a tree may not provide a complete picture of their history. In such cases,tree-based phylogenetic networks can provide a useful, more refined representation of the species evolution. Such a network is essentially a phylogenetic tree with some arcs added between the tree edges so as to represent reticulate events such as gene transfer. Even so, this model does not permit the representation of evolutionary scenarios where reticulate events have taken place between different subfamilies or lineages of species. To represent such scenarios, in this paper we introduce the notion of a forest-based phylogenetic network, that is, a collection of leaf-disjoint phylogenetic trees on a set of species with arcs added between the edges of distinct trees within the collection. Forest-based networks include the recently introduced class of overlaid species forests which are used to model introgression. As we shall see, even though the definition of forest-based networks is closely related to that of tree-based networks, they lead to new mathematical theory which complements that of tree-based networks. As well as studying the relationship of forest-based networks with other classes of phylogenetic networks, such as tree-child networks and universal tree-based networks, we present some characterizations of some special classes of forest-based networks. We expect that our results will be useful for developing new models and algorithms to understand reticulate evolution, such as gene transfer between collections of bacteria that live in different environments.
What makes a reaction network "chemical"?
Reaction networks (RNs) comprise a set $X$ of species and a set $\mathscr{R}$ of reactions $Y\to Y'$, each converting a multiset of educts $Y\subseteq X$ into a multiset $Y'\subseteq X$ of products. RNs are equivalent to directed hypergraphs. However, not all RNs necessarily admit a chemical interpretation. Instead, they might contradict fundamental principles of physics such as the conservation of energy and mass or the reversibility of chemical reactions. The consequences of these necessary conditions for the stoichiometric matrix $\mathbf{S} \in \mathbb{R}^{X\times\mathscr{R}}$ have been discussed extensively in the literature. Here, we provide sufficient conditions for $\mathbf{S}$ that guarantee the interpretation of RNs in terms of balanced sum formulas and structural formulas, respectively. Chemically plausible RNs allow neither a perpetuum mobile, i.e., a "futile cycle" of reactions with non-vanishing energy production, nor the creation or annihilation of mass. Such RNs are said to be thermodynamically sound and conservative. For finite RNs, both conditions can be expressed equivalently as properties of $\mathbf{S}$. The first condition is vacuous for reversible networks, but it excludes irreversible futile cycles and - in a stricter sense - futile cycles that even contain an irreversible reaction. The second condition is equivalent to the existence of a strictly positive reaction invariant. Furthermore, it is sufficient for the existence of a realization in terms of sum formulas, obeying conservation of "atoms". In particular, these realizations can be chosen such that any two species have distinct sum formulas, unless $\mathbf{S}$ implies that they are "obligatory isomers". In terms of structural formulas, every compound is a labeled multigraph, in essence a Lewis formula, and reactions comprise only a rearrangement of bonds such that the total bond order is preserved.
2021-12-31 v5
Introducing DASEP: the doubly asymmetric simple exclusion process
Published in "Séminaire Lotharingien de Combinatoire" 87B (2023), pp. 81-91 • Search Publication
Research in combinatorics has often explored the asymmetric simple exclusion process (ASEP). The ASEP, inspired by examples from statistical mechanics, involves particles of various species moving around a lattice. With the traditional ASEP particles of a given species can move but do not change species. In this paper a new combinatorial formalism, the DASEP (doubly asymmetric simple exclusion process), is explored. The DASEP is inspired by biological processes where, unlike the ASEP, the particles can change from one species to another. The combinatorics of the DASEP on a one dimensional lattice are explored, including the associated generating function. The stationary probabilities of the DASEP are explored, and results are proven relating these stationary probabilities to those of the simpler ASEP.
2021-11-25
On the quartet distance given partial information
Published • View Publication • BIB
Let $T$ be an arbitrary phylogenetic tree with $n$ leaves. It is well-known that the average quartet distance between two assignments of taxa to the leaves of $T$ is $\frac 23 \binom{n}{4}$. However, a longstanding conjecture of Bandelt and Dress asserts that $(\frac 23 +o(1))\binom{n}{4}$ is also the {\em maximum} quartet distance between two assignments. While Alon, Naves, and Sudakov have shown this indeed holds for caterpillar trees, the general case of the conjecture is still unresolved. A natural extension is when partial information is given: the two assignments are known to coincide on a given subset of taxa. The partial information setting is biologically relevant as the location of some taxa (species) in the phylogenetic tree may be known, and for other taxa it might not be known. What can we then say about the average and maximum quartet distance in this more general setting? Surprisingly, even determining the {\em average} quartet distance becomes a nontrivial task in the partial information setting and determining the maximum quartet distance is even more challenging, as these turn out to be dependent of the structure of $T$. In this paper we prove nontrivial asymptotic bounds that are sometimes tight for the average quartet distance in the partial information setting. We also show that the Bandelt and Dress conjecture does not generally hold under the partial information setting. Specifically, we prove that there are cases where the average and maximum quartet distance substantially differ.
2021-11-24
Convex characters, algorithms and matchings
Published • View Publication • BIB
Phylogenetic trees are used to model evolution: leaves are labelled to represent contemporary species ("taxa") and interior vertices represent extinct ancestors. Informally, convex characters are measurements on the contemporary species in which the subset of species (both contemporary and extinct) that share a given state, form a connected subtree. In \cite{KelkS17} it was shown how to efficiently count, list and sample certain restricted subfamilies of convex characters, and algorithmic applications were given. We continue this work in a number of directions. First, we show how combining the enumeration of convex characters with existing parameterised algorithms can be used to speed up exponential-time algorithms for the \emph{maximum agreement forest problem} in phylogenetics. Second, we re-visit the quantity $g_2(T)$, defined as the number of convex characters on $T$ in which each state appears on at least 2 taxa. We use this to give an algorithm with running time $O( φ^{n} \cdot \text{poly}(n) )$, where $φ\approx 1.6181$ is the golden ratio and $n$ is the number of taxa in the input trees, for computation of \emph{maximum parsimony distance on two state characters}. By further restricting the characters counted by $g_2(T)$ we open an interesting bridge to the literature on enumeration of matchings. By crossing this bridge we improve the running time of the aforementioned parsimony distance algorithm to $O( 1.5895^{n} \cdot \text{poly}(n) )$, and obtain a number of new results in themselves relevant to enumeration of matchings on at-most binary trees.
2021-11-19 v3
A lattice structure for ancestral configurations arising from the relationship between gene trees and species trees
Published • View Publication • BIB
To a given gene tree topology $G$ and species tree topology $S$ with leaves labeled bijectively from a fixed set $X$, one can associate a set of ancestral configurations, each of which encodes a set of gene lineages that can be found at a given node of the species tree. We introduce a lattice structure on ancestral configurations, studying the directed graphs that provide graphical representations of lattices of ancestral configurations. For a matching gene tree topology and species tree topology $G=S$, we present a method for defining the digraph of ancestral configurations from the tree topology by using iterated cartesian products of graphs. We show that a specific set of paths on the digraph of ancestral configurations is in bijection with the set of labeled histories -- a well-known phylogenetic object that enumerates possible temporal orderings of the coalescences of a tree. For each of a series of tree families, we obtain closed-form expressions for the number of labeled histories by using this bijection to count paths on associated digraphs. Finally, we prove that our lattice construction extends to nonmatching tree pairs, and we use it to characterize pairs $(G,S)$ having the maximal number of ancestral configurations for a fixed $G$. We discuss how the construction provides new methods for performing enumerations of combinatorial aspects of gene and species trees.
Quasi-Best Match Graphs
Quasi-best match graphs (qBMGs) are a hereditary class of directed, properly vertex-colored graphs. They arise naturally in mathematical phylogenetics as a generalization of best match graphs, which formalize the notion of evolutionary closest relatedness of genes (vertices) in multiple species (vertex colors). They are explained by rooted trees whose leaves correspond to vertices. In contrast to BMGs, qBMGs represent only best matches at a restricted phylogenetic distance. We provide characterizations of qBMGs that give rise to polynomial-time recognition algorithms and identify the BMGs as the qBMGs that are color-sink-free. Furthermore, two-colored qBMGs are characterized as directed graphs satisfying three simple local conditions, two of which have appeared previously, namely bi-transitivity in the sense of Das et al. (2021) and a hierarchy-like structure of out-neighborhoods, i.e., $N(x)\cap N(y)\in\{N(x),N(y),\emptyset\}$ for any two vertices $x$ and $y$. Further results characterize qBMGs that can be explained by binary phylogenetic trees.
Encoding and ordering X-cactuses
Published • View Publication • BIB
Phylogenetic networks are a generalization of evolutionary or phylogenetic trees that are commonly used to represent the evolution of species which cross with one another. A special type of phylogenetic network is an {\em $X$-cactus}, which is essentially a cactus graph in which all vertices with degree less than three are labelled by at least one element from a set $X$ of species. In this paper, we present a way to {\em encode} $X$-cactuses in terms of certain collections of partitions of $X$ that naturally arise from $X$-cactuses. Using this encoding, we also introduce a partial order on the set of $X$-cactuses (up to isomorphism), and derive some structural properties of the resulting partially ordered set. This includes an analysis of some properties of its least upper and greatest lower bounds. Our results not only extend some fundamental properties of phylogenetic trees to $X$-cactuses, but also provides a new approach to solving topical problems in phylogenetic network theory such as deriving consensus networks.
2021-07-22 v3
"Normal" phylogenetic networks may be emerging as the leading class
Published • View Publication • BIB
The rich and varied ways that genetic material can be passed between species has motivated extensive research into the theory of phylogenetic networks. Features that align with biological processes, or with desirable mathematical properties, have been used to define classes and prove results, with the goal of developing the theoretical foundations for network reconstruction methods. We may have now reached the point where a collection of recent results can be drawn together to make one class of network, the \emph{normal} networks, a leading contender, sitting in the sweet spot between biological relevance and mathematical tractability.
On the Complexity of Optimising Variants of Phylogenetic Diversity on Phylogenetic Networks
Published • View Publication • BIB
Phylogenetic Diversity (PD) is a prominent quantitative measure of the biodiversity of a collection of present-day species (taxa). This measure is based on the evolutionary distance among the species in the collection. Loosely speaking, if $\mathcal{T}$ is a rooted phylogenetic tree whose leaf set $X$ represents a set of species and whose edges have real-valued lengths (weights), then the PD score of a subset $S$ of $X$ is the sum of the weights of the edges of the minimal subtree of $\mathcal{T}$ connecting the species in $S$. In this paper, we define several natural variants of the PD score for a subset of taxa which are related by a known rooted phylogenetic network. Under these variants, we explore, for a positive integer $k$, the computational complexity of determining the maximum PD score over all subsets of taxa of size $k$ when the input is restricted to different classes of rooted phylogenetic networks