species
269 papers tagged with this keyword
A species approach to Rota's twelvefold way
An introduction to Joyal's theory of combinatorial species is given and through it an alternative view of Rota's twelvefold way emerges.
Reciprocal Best Match Graphs
Reciprocal best matches play an important role in numerous applications in computational biology, in particular as the basis of many widely used tools for orthology assessment. Nevertheless, very little is known about their mathematical structure. Here, we investigate the structure of reciprocal best match graphs (RBMGs). In order to abstract from the details of measuring distances, we define reciprocal best matches here as pairwise most closely related leaves in a gene tree, arguing that conceptually this is the notion that is pragmatically approximated by distance- or similarity-based heuristics. We start by showing that a graph $G$ is an RBMG if and only if its quotient graph w.r.t.\ a certain thinness relation is an RBMG. Furthermore, it is necessary and sufficient that all connected components of $G$ are RBMGs. The main result of this contribution is a complete characterization of RBMGs with 3 colors/species that can be checked in polynomial time. For 3 colors, there are three distinct classes of trees that are related to the structure of the phylogenetic trees explaining them. We derive an approach to recognize RBMGs with an arbitrary number of colors; it remains open however, whether a polynomial-time for RBMG recognition exists. In addition, we show that RBMGs that at the same time are cographs (co-RBMGs) can be recognized in polynomial time. Co-RBMGs are characterized in terms of hierarchically colored cographs, a particular class of vertex colored cographs that is introduced here. The (least resolved) trees that explain co-RBMGs can be constructed in polynomial time.
Hereditary species as monoidal decomposition spaces, comodule bialgebras, and operadic categories
We show that Schmitt's hereditary species induce monoidal decomposition spaces, and exhibit Schmitt's bialgebra construction as an instance of the general bialgebra construction on a monoidal decomposition space. We show furthermore that this bialgebra structure coacts on the underlying restriction-species bialgebra structure so as to form a comodule bialgebra. Finally, we show that hereditary species induce a new family of examples of operadic categories in the sense of Batanin and Markl.
Combinatorial properties of phylogenetic diversity indices
Phylogenetic diversity indices provide a formal way to apportion 'evolutionary heritage' across species. Two natural diversity indices are Fair Proportion (FP) and Equal Splits (ES). FP is also called 'evolutionary distinctiveness' and, for rooted trees, is identical to the Shapley Value (SV), which arises from cooperative game theory. In this paper, we investigate the extent to which FP and ES can differ, characterise tree shapes on which the indices are identical, and study the equivalence of FP and SV and its implications in more detail. We also define and investigate analogues of these indices on unrooted trees (where SV was originally defined), including an index that is closely related to the Pauplin representation of phylogenetic diversity.
The exact phase diagram for a semipermeable TASEP with nonlocal boundary jumps
Published in J. Phys. A: Math. Theor. 52 (2019) 355001 (19pp)
• View Publication
• BIB
We consider a finite one-dimensional totally asymmetric simple exclusion process (TASEP) with four types of particles, $\{1,0,\bar{1},*\}$, in contact with reservoirs. Particles of species $0$ can neither enter nor exit the lattice, and those of species $*$ are constrained to lie at the first and last site. Particles of species $1$ enter from the left reservoir into either the first or second site, move rightwards, and leave from either the last or penultimate site. Conversely, particles of species $\bar{1}$ enter from the right reservoir into either the last or penultimate site, move leftwards, and leave from either the first or last site. This dynamics is motivated by a natural random walk on the Weyl group of type D. We compute the exact nonequilibrium steady state distribution using a matrix ansatz building on earlier work of Arita. We then give explicit formulas for the nonequilibrium partition function as well as densities and currents of all species in the steady state, and derive the phase diagram.
Displaying trees across two phylogenetic networks
Published in Theoretical Computer Science, 796:129-146, 2020
• View Publication
• BIB
Phylogenetic networks are a generalization of phylogenetic trees to leaf-labeled directed acyclic graphs that represent ancestral relationships between species whose past includes non-tree-like events such as hybridization and horizontal gene transfer. Indeed, each phylogenetic network embeds a collection of phylogenetic trees. Referring to the collection of trees that a given phylogenetic network $N$ embeds as the display set of $N$, several questions in the context of the display set of $N$ have recently been analyzed. For example, the widely studied Tree-Containment problem asks if a given phylogenetic tree is contained in the display set of a given network. The focus of this paper are two questions that naturally arise in comparing the display sets of two phylogenetic networks. First, we analyze the problem of deciding if the display sets of two phylogenetic networks have a tree in common. Surprisingly, this problem turns out to be NP-complete even for two temporal normal networks. Second, we investigate the question of whether or not the display sets of two phylogenetic networks are equal. While we recently showed that this problem is polynomial-time solvable for a normal and a tree-child network, it is computationally hard in the general case. In establishing hardness, we show that the problem is contained in the second level of the polynomial-time hierarchy. Specifically, it is $Π_2^P$-complete. Along the way, we show that two other problems are also $Π_2^P$-complete, one of which being a generalization of Tree-Containment.
A class of phylogenetic networks reconstructable from ancestral profiles
Published
• View Publication
• BIB
Rooted phylogenetic networks provide an explicit representation of the evolutionary history of a set $X$ of sampled species. In contrast to phylogenetic trees which show only speciation events, networks can also accommodate reticulate processes (for example, hybrid evolution, endosymbiosis, and lateral gene transfer). A major goal in systematic biology is to infer evolutionary relationships, and while phylogenetic trees can be uniquely determined from various simple combinatorial data on $X$, for networks the reconstruction question is much more subtle. Here we ask when can a network be uniquely reconstructed from its `ancestral profile' (the number of paths from each ancestral vertex to each element in $X$). We show that reconstruction holds (even within the class of all networks) for a class of networks we call `orchard networks', and we provide a polynomial-time algorithm for reconstructing any orchard network from its ancestral profile. Our approach relies on establishing a structural theorem for orchard networks, which also provides for a fast (polynomial-time) algorithm to test if any given network is of orchard type. Since the class of orchard networks includes tree-sibling tree-consistent networks and tree-child networks, our result generalise reconstruction results from 2008 and 2009. Orchard networks allow for an unbounded number $k$ of reticulation vertices, in contrast to tree-sibling tree-consistent networks and tree-child networks for which $k$ is at most $2|X|-4$ and $|X|-1$, respectively.
The adjoint braid arrangement as a combinatorial Lie algebra via the Steinmann relations
We study a certain discrete differentiation of piecewise-constant functions on the adjoint of the braid hyperplane arrangement, defined by taking finite-differences across hyperplanes. In terms of Aguiar-Mahajan's Lie theory of hyperplane arrangements, we show that this structure is equivalent to the action of Lie elements on faces. We use layered binary trees to encode flags of adjoint arrangement faces, allowing for the representation of certain Lie elements by antisymmetrized layered binary forests. This is dual to the well-known use of (delayered) binary trees to represent Lie elements of the braid arrangement. The discrete derivative then induces an action of layered binary forests on piecewise-constant functions, which we call the forest derivative. Our main result states that forest derivatives of functions factorize as external products of functions precisely if one restricts to functions which satisfy the Steinmann relations, which are certain four-term linear relations appearing in the foundations of axiomatic quantum field theory. We also show that the forest derivative satisfies the Lie properties of antisymmetry the Jacobi identity. It follows from these Lie properties, and also crucially factorization, that functions which satisfy the Steinmann relations form a left comodule of the Lie cooperad, with the coaction given by the forest derivative. Dually, this endows the adjoint braid arrangement modulo the Steinmann relations with the structure of a Lie algebra internal to the category of vector species. This work is a first step towards describing new connections between Hopf theory in species and quantum field theory.
Möbius functions of directed restriction species and free operads, via the generalised Rota formula
We present some tools for providing situations where the generalised Rota formula of arXiv:1801.07504 applies. As an example of this, we compute the Möbius function of the incidence algebra of any directed restriction species, free operad, or more generally free monad on a finitary polynomial monad.
Reconciling Event-Labeled Gene Trees with MUL-trees and Species Networks
Phylogenomics commonly aims to construct evolutionary trees from genomic sequence information. One way to approach this problem is to first estimate event-labeled gene trees (i.e., rooted trees whose non-leaf vertices are labeled by speciation or gene duplication events), and to then look for a species tree which can be reconciled with this tree through a \emph{reconciliation map} between the trees. In practice, however, it can happen that there is no such map from a given event-labeled tree to \emph{any} species tree. An important situation where this might arise is where the species evolution is better represented by a \emph{network} instead of a tree. In this paper, we therefore consider the problem of reconciling event-labeled trees with species networks. In particular, we prove that any event-labeled gene tree can be reconciled with some network and that, under certain mild assumptions on the gene tree, the network can even be assumed to be multi-arc free. To prove this result, we show that we can always reconcile the gene tree with some multi-labeled (MUL-)tree, which can then be "folded up" to produce the desired reconciliation and network. In addition, we study the interplay between reconciliation maps from event-labeled gene trees to MUL-trees and networks. Our results could be useful for understanding how genomes have evolved after undergoing complex evolutionary events such as polyploidy.
Quantifying CDS Sortability of Permutations by Strategic Pile Size
The special purpose sorting operation, context directed swap (CDS), is an example of the block interchange sorting operation studied in prior work on permutation sorting. CDS has been postulated to model certain molecular sorting events that occur in the genome maintenance program of some species of ciliates. We investigate the mathematical structure of permutations not sortable by the CDS sorting operation. In particular, we present substantial progress towards quantifying permutations with a given strategic pile size, which can be understood as a measure of CDS non-sortability. Our main results include formulas for the number of permutations in $\textsf{S}_n$ with maximum size strategic pile. More generally, we derive a formula for the number of permutations in $\textsf{S}_n$ with strategic pile size $k$, in addition to an algorithm for computing certain coefficients of this formula, which we call merge numbers.
On chordal phylogeny graphs
Published
• View Publication
• BIB
An acyclic digraph each vertex of which has indegree at most $i$ and outdegree at most $j$ is called an $(i, j)$ digraph for some positive integers $i$ and $j$. Lee {\it et al.} (2017) studied the phylogeny graphs of $(2, 2)$ digraphs and gave sufficient conditions and necessary conditions for $(2, 2)$ digraphs having chordal phylogeny graphs. Their work was motivated by problems related to evidence propagation in a Bayesian network for which it is useful to know which acyclic digraphs have their moral graphs being chordal (phylogeny graphs are called moral graphs in Bayesian network theory).
In this paper, we extend their work. We completely characterize phylogeny graphs of $(1, i)$ digraphs and $(i,1)$ digraphs, respectively, for a positive integer $i$. Then, we study phylogeny graphs of a $(2,j)$ digraphs, which is worthwhile in the context that a child has two biological parents in most species, to show that the phylogeny graph of a $(2,j)$ digraph $D$ is chordal if the underlying graph of $D$ is chordal for any positive integer $j$. Especially, we show that as long as the underlying graph of a $(2,2)$ digraph is chordal, its phylogeny graph is not only chordal but also planar.
Ranked Schröder Trees
Published
• View Publication
• BIB
In biology, a phylogenetic tree is a tool to represent the evolutionary relationship between species. Unfortunately, the classical Schröder tree model is not adapted to take into account the chronology between the branching nodes. In particular, it does not answer the question: how many different phylogenetic stories lead to the creation of n species and what is the average time to get there? In this paper, we enrich this model in two distinct ways in order to obtain two ranked tree models for phylogenetics, i.e. models coding chronology. For that purpose, we first develop a model of (strongly) increasing Schröder trees, symbolically described in the classical context of increasing labeling. Then we introduce a generalization for the labeling with some unusual order constraint in Analytic Combinatorics (namely the weakly increasing trees). Although these models are direct extensions of the Schröder tree model, it appears that they are also in one-to-one correspondence with several classical combinatorial objects. Through the paper, we present these links, exhibit some parameters in typical large trees and conclude the studies with efficient uniform samplers.
On the uniqueness of the maximum parsimony tree for data with up to two substitutions: an extension of the classic Buneman theorem in phylogenetics
Published
• View Publication
• BIB
One of the main aims of phylogenetics is the reconstruction of the correct evolutionary tree when data concerning the underlying species set are given. These data typically come in the form of DNA, RNA or protein alignments, which consist of various characters (also often referred to as sites). Often, however, tree reconstruction methods based on criteria like maximum parsimony may fail to provide a unique tree for a given dataset, or, even worse, reconstruct the `wrong' tree (i.e. a tree that differs from the one that generated the data). On the other hand it has long been known that if the alignment consists of all the characters that correspond to edges of a particular tree, i.e. they all require exactly $k=1$ substitution to be realized on that tree, then this tree will be recovered by maximum parsimony methods. This is based on Buneman's theorem in mathematical phylogenetics. It is the goal of the present manuscript to extend this classic result as follows: We prove that if an alignment consists of all characters that require exactly $k=2$ substitutions on a particular tree, this tree will always be the unique maximum parsimony tree (and we also show that this can be generalized to characters which require at most $k=2$ substitutions). In particular, this also proves a conjecture based on a recently published observation by Goloboff et al. affirmatively for the special case of $k=2$.
On the Subnet Prune and Regraft Distance
Published in The Electronic Journal of Combinatorics 26(2) (2019), #P2.3
• View Publication
• BIB
Phylogenetic networks are rooted directed acyclic graphs that represent evolutionary relationships between species whose past includes reticulation events such as hybridisation and horizontal gene transfer. To search the space of phylogenetic networks, the popular tree rearrangement operation rooted subtree prune and regraft (rSPR) was recently generalised to phylogenetic networks. This new operation - called subnet prune and regraft (SNPR) - induces a metric on the space of all phylogenetic networks as well as on several widely-used network classes. In this paper, we investigate several problems that arise in the context of computing the SNPR-distance. For a phylogenetic tree $T$ and a phylogenetic network $N$, we show how this distance can be computed by considering the set of trees that are embedded in $N$ and then use this result to characterise the SNPR-distance between $T$ and $N$ in terms of agreement forests. Furthermore, we analyse properties of shortest SNPR-sequences between two phylogenetic networks $N$ and $N'$, and answer the question whether or not any of the classes of tree-child, reticulation-visible, or tree-based networks isometrically embeds into the class of all phylogenetic networks under SNPR.
On integral structure types
Published
• View Publication
• BIB
We introduce integral structure types as a categorical analogue of virtual combinatorial species. Integral structure types then categorify power series with possibly negative coefficients in the same way that combinatorial species categorify power series with non-negative rational coefficients. The notion of an operator on combinatorial species naturally extends to integral structure types, and in light of their `negativity' we define the notion of the commutator of two operators on integral structure types. We then extend integral structure types to the setting of stuff types as introduced by Baez and Dolan, and then conclude by using integral structure types to give a combinatorial description for Chern classes of projective hypersurfaces.
Eulerian polynomials on segmented permutations
We define a generalization of the Eulerian polynomials and the Eulerian numbers by considering a descent statistic on segmented permutations coming from the study of 2-species exclusion processes and a change of basis in a Hopf algebra. We give some properties satisfied by these generalized Eulerian numbers. We also define a $q$-analog of these Eulerian polynomials which gives back usual Eulerian polynomials and ordered Bell polynomials for specific values of its variables. We also define a noncommutative analog living in the algebra of segmented compositions. It gives us an explicit generating function and some identities satisfied by the generalized Eulerian polynomials such as a Worpitzky-type relation.
Phylogenetic networks that are their own fold-ups
Published
• View Publication
• BIB
Phylogenetic networks are becoming of increasing interest to evolutionary biologists due to their ability to capture complex non-treelike evolutionary processes. From a combinatorial point of view, such networks are certain types of rooted directed acyclic graphs whose leaves are labelled by, for example, species. A number of mathematically interesting classes of phylogenetic networks are known. These include the biologically relevant class of stable phylogenetic networks whose members are defined via certain "fold-up" and "un-fold" operations that link them with concepts arising within the theory of, for example, graph fibrations. Despite this exciting link, the structural complexity of stable phylogenetic networks is still relatively poorly understood. Employing the popular tree-based, reticulation-visible, and tree-child properties which allow one to gauge this complexity in one way or another, we provide novel characterizations for when a stable phylogenetic network satisfies either one of these three properties.
Best Match Graphs
Published
• View Publication
• BIB
THIS IS A CORRECTED VERSION INCLUDING AN APPENDED CORRIGENDUM.
Best match graphs arise naturally as the first processing intermediate in algorithms for orthology detection. Let $T$ be a phylogenetic (gene) tree $T$ and $σ$ an assignment of leaves of $T$ to species. The best match graph $(G,σ)$ is a digraph that contains an arc from $x$ to $y$ if the genes $x$ and $y$ reside in different species and $y$ is one of possibly many (evolutionary) closest relatives of $x$ compared to all other genes contained in the species $σ(y)$. Here, we characterize best match graphs and show that it can be decided in cubic time and quadratic space whether $(G,σ)$ derived from a tree in this manner. If the answer is affirmative, there is a unique least resolved tree that explains $(G,σ)$, which can also be constructed in cubic time.
Split graphs: combinatorial species and asymptotics
Published in Electron. J. Combin. 26 (2019), #P2.42
• View Publication
• BIB
A split graph is a graph whose vertices can be partitioned into a clique and a stable set. We investigate the combinatorial species of split graphs, providing species-theoretic generalizations of enumerative results due to Bína and Přibil (2015), Cheng, Collins, and Trenk (2016), and Collins and Trenk (2018). In both the labeled and unlabeled cases, we give asymptotic results on the number of split graphs, of unbalanced split graphs, and of bicolored graphs, including proving the conjecture of Cheng, Collins, and Trenk (2016) that almost all split graphs are balanced.