species
269 papers tagged with this keyword
A complete characterization of pairs of binary phylogenetic trees with identical $A_k$-alignments
Published
• View Publication
• BIB
Phylogenetic trees play a key role in the reconstruction of evolutionary relationships. Typically, they are derived from aligned sequence data (like DNA, RNA, or proteins) by using optimization criteria like, e.g., maximum parsimony (MP). It is believed that the latter is able to reconstruct the \enquote{true} tree, i.e., the tree that generated the data, whenever the number of substitutions required to explain the data with that tree is relatively small compared to the size of the tree (measured in the number $n$ of leaves of the tree, which represent the species under investigation). However, reconstructing the correct tree from any alignment first and foremost requires the given alignment to perform differently on the \enquote{correct} tree than on others.
A special type of alignments, namely so-called $A_k$-alignments, has gained considerable interest in recent literature. These alignments consist of all binary characters (\enquote{sites}) which require precisely $k$ substitutions on a given tree. It has been found that whenever $k$ is small enough (in comparison to $n$), $A_k$-alignments uniquely characterize the trees that generated them. However, recent literature has left a significant gap between $n\leq 2k+2$ -- namely the cases in which no such characterization is possible -- and $n\geq 4k$ -- namely the cases in which this characterization works. It is the main aim of the present manuscript to close this gap, i.e., to present a full characterization of all pairs of trees that share the same $A_k$-alignment. In particular, we show that indeed every binary phylogenetic tree with $n$ leaves is uniquely defined by its $A_k$-alignments if $n\geq 2k+3$. By closing said gap, we also ensure that our result is optimal.
Compactifications of phylogenetic systems and species of electrical networks
We describe new spaces and maps. Our graphical map is a visual and numerical correspondence between spaces of circular electrical networks and circular planar split systems. When restricted to the planar circular electrical case, this graphical map finds the split system uniquely associated with the Kalmanson resistance distance of the dual network, matching the induced split system familiar from phylogenetics. This correspondence is extended to compactifications of the respective spaces, taking cactus networks to the cactus split systems defined herein. The graphical map preserves both network components and cactus structure, allowing an elegant enumeration of induced phylogenetic split systems via combinatorial species. We introduce the global spaces of circular planar electrical networks and circular split systems. These new spaces are also CW complexes, but the 0-cells of each are counted by the Bell numbers as opposed to the Catalan numbers. As species, the two sorts of global cacti are seen to be compositions in complementary ways.
Pattern-avoiding Cayley permutations via combinatorial species
A Cayley permutation is a word of positive integers such that if a letter appears in this word, then all positive integers smaller than that letter also appear. We initiate a systematic study of pattern avoidance on Cayley permutations adopting a combinatorial species approach. Our methods lead to species equations, generating series, and counting formulas for Cayley permutations avoiding any pattern of length at most three. We also introduce the species of primitive structures as a generalization of Cayley permutations with no "flat steps". Finally, we explore various notions of Wilf equivalence arising in this context.
Colouring isonemal fabrics with more than two colours and non-twilly redundancy
Perfect colouring of isonemal fabrics by thin and thick striping of warp and weft with more than two colours is examined where the cells with warps and wefts of the same colour do not appear along diagonal lines (not twilly redundancy). In species 33 to 39, all fabrics of specific orders (linear periods) can be perfectly coloured by thin striping with satin redundancy. In fewer species, all fabrics of specific orders can be perfectly coloured by thick striping in two different ways with doubled satin redundancy. The isonemal fabric 6-1-1 can be used as a redundancy configuration. All these colourings allow the coloured weaving of flat tori and -- a few of them -- cubes.
Bounding the softwired parsimony score of a phylogenetic network
In comparison to phylogenetic trees, phylogenetic networks are more suitable to represent complex evolutionary histories of species whose past includes reticulation such as hybridisation or lateral gene transfer. However, the reconstruction of phylogenetic networks remains challenging and computationally expensive due to their intricate structural properties. For example, the small parsimony problem that is solvable in polynomial time for phylogenetic trees, becomes NP-hard on phylogenetic networks under softwired and parental parsimony, even for a single binary character and structurally constrained networks. To calculate the parsimony score of a phylogenetic network $N$, these two parsimony notions consider different exponential-size sets of phylogenetic trees that can be extracted from $N$ and infer the minimum parsimony score over all trees in the set. In this paper, we ask: What is the maximum difference between the parsimony score of any phylogenetic tree that is contained in the set of considered trees and a phylogenetic tree whose parsimony score equates to the parsimony score of $N$? Given a gap-free sequence alignment of multi-state characters and a rooted binary level-$k$ phylogenetic network, we use the novel concept of an informative blob to show that this difference is bounded by $k+1$ times the softwired parsimony score of $N$. In particular, the difference is independent of the alignment length and the number of character states. We show that an analogous bound can be obtained for the softwired parsimony score of semi-directed networks, while under parental parsimony on the other hand, such a bound does not hold.
A Vector Representation for Phylogenetic Trees
Published
• View Publication
• BIB
Good representations for phylogenetic trees and networks are important for optimizing storage efficiency and implementation of scalable methods for the inference and analysis of evolutionary trees for genes, genomes and species. We introduce a new representation for rooted phylogenetic trees that encodes a binary tree on n taxa as a vector of length 2n in which each taxon appears exactly twice. Using this new tree representation, we introduce a novel tree rearrangement operator, called a HOP, that results in a tree space of diameter n and a quadratic neighbourhood size. We also introduce a novel metric, the HOP distance, which is the minimum number of HOPs to transform a tree into another tree. The HOP distance can be computed in near-linear time, a rare instance of a tree rearrangement distance that is tractable. Our experiments show that the HOP distance is better correlated to the Subtree-Prune-and-Regraft distance than the widely used Robinson-Foulds distance. We also describe how the novel tree representation we introduce can be further generalized to tree-child networks.
Operad structures on the species composition of two operads
Published
• View Publication
• BIB
We give an explicit description of three operad structures on the species composition $p \circ q$, where $q$ is any given positive operad, and where $p$ is the NAP operad, or a shuffle version of the magmatic operad Mag. No distributive law between $p$ and $q$ is assumed.
Phylogenetic diversity indices from an affine and projective viewpoint
Published
• View Publication
• BIB
Phylogenetic diversity indices are commonly used to rank the elements in a collection of species or populations for conservation purposes. The derivation of these indices is typically based on some quantitative description of the evolutionary history of the species in question, which is often given in terms of a phylogenetic tree. Both rooted and unrooted phylogenetic trees can be employed, and there are close connections between the indices that are derived in these two different ways. In this paper, we introduce more general phylogenetic diversity indices that can be derived from collections of subsets (clusters) and collections of bipartitions (splits) of the given set of species. Such indices could be useful, for example, in case there is some uncertainty in the topology of the tree being used to derive a phylogenetic diversity index. As well as characterizing some of the indices that we introduce in terms of their special properties, we provide a link between cluster-based and split-based phylogenetic diversity indices that uses a discrete analogue of the classical link between affine and projective geometry. This provides a unified framework for many of the various phylogenetic diversity indices used in the literature based on rooted and unrooted phylogenetic trees, generalizations and new proofs for previous results concerning tree-based indices, and a way to define some new phylogenetic diversity indices that naturally arise as affine or projective variants of each other.
Exact and Heuristic Computation of the Scanwidth of Directed Acyclic Graphs
Published
• View Publication
• BIB
To measure the tree-likeness of a directed acyclic graph (DAG), a new width parameter that considers the directions of the arcs was recently introduced: scanwidth. We present the first algorithm that efficiently computes the exact scanwidth of general DAGs. For DAGs with one root and scanwidth $k$ it runs in $O(k \cdot n^k \cdot m)$ time. The algorithm also functions as an FPT algorithm with complexity $O(2^{4 \ell - 1} \cdot \ell \cdot n + n^2)$ for phylogenetic networks of level-$\ell$, a type of DAG used to depict evolutionary relationships among species. Our algorithm performs well in practice, being able to compute the scanwidth of synthetic networks up to 30 reticulations and 100 leaves within 500 seconds. Furthermore, we propose a heuristic that obtains an average practical approximation ratio of 1.5 on these networks. While we prove that the scanwidth is bounded from below by the treewidth of the underlying undirected graph, experiments suggest that for networks the parameters are close in practice.
The inhomogeneous $t$-PushTASEP and Macdonald polynomials
Published
• View Publication
• BIB
We study a multispecies $t$-PushTASEP system on a finite ring of $n$ sites with site-dependent rates $x_1,\dots,x_n$. Let $λ=(λ_1,\dots,λ_n)$ be a partition whose parts represent the species of the $n$ particles on the ring. We show that for each composition $η$ obtained by permuting the parts of $λ$, the stationary probability of being in state $η$ is proportional to the ASEP polynomial $F_η(x_1,\dots,x_n; q,t)$ at $q=1$; the normalizing constant (or partition function) is the Macdonald polynomial $P_λ(x_1,\dots,x_n;q,t)$ at $q=1$. Our approach involves new relations between the families of ASEP polynomials and of non-symmetric Macdonald polynomials at $q=1$. We also use multiline diagrams, showing that a single jump of the PushTASEP system is closely related to the operation of moving from one line to the next in a multiline diagram. We derive symmetry properties for the system under permutation of its jump rates, as well as a formula for the current of a single-species system.
On the correctness of Maximum Parsimony for data with few substitutions in the NNI neighborhood of phylogenetic trees
Published
• View Publication
• BIB
Estimating phylogenetic trees, which depict the relationships between different species, from aligned sequence data (such as DNA, RNA, or proteins) is one of the main aims of evolutionary biology. However, tree reconstruction criteria like maximum parsimony do not necessarily lead to unique trees and in some cases even fail to recognize the \enquote{correct} tree (i.e., the tree on which the data was generated). On the other hand, a recent study has shown that for an alignment containing precisely those binary characters (sites) which require up to two substitutions on a given tree, this tree will be the unique maximum parsimony tree.
It is the aim of the present paper to generalize this recent result in the following sense: We show that for a tree $T$ with $n$ leaves, as long as $k<\frac{n}{8}+\frac{11}{9}-\frac{1}{18}\sqrt{9\cdot \left(\frac{n}{4}\right)^2+16}$ (or, equivalently, $n>9 k-11+\sqrt{9k^2-22 k+17} $, which in particular holds for all $n\geq 12k$), the maximum parsimony tree for the alignment containing all binary characters which require (up to or precisely) $k$ substitutions on $T$ will be unique in the NNI neighborhood of $T$ and it will coincide with $T$, too. In other words, within the NNI neighborhood of $T$, $T$ is the unique most parsimonious tree for the said alignment. This partially answers a recently published conjecture affirmatively. Additionally, we show that for $n\geq 8$ and for $k$ being in the order of $\frac{n}{2}$, there is always a pair of phylogenetic trees $T$ and $T'$ which are NNI neighbors, but for which the alignment of characters requiring precisely $k$ substitutions each on $T$ in total requires fewer substitutions on $T'$.
Identifying circular orders for blobs in phylogenetic networks
Published
• View Publication
• BIB
Interest in the inference of evolutionary networks relating species or populations has grown with the increasing recognition of the importance of hybridization, gene flow and admixture, and the availability of large-scale genomic data. However, what network features may be validly inferred from various data types under different models remains poorly understood. Previous work has largely focused on level-1 networks, in which reticulation events are well separated, and on a general network's tree of blobs, the tree obtained by contracting every blob to a node. An open question is the identifiability of the topology of a blob of unknown level. We consider the identifiability of the circular order in which subnetworks attach to a blob, first proving that this order is well-defined for outer-labeled planar blobs. For this class of blobs, we show that the circular order information from 4-taxon subnetworks identifies the full circular order of the blob. Similarly, the circular order from 3-taxon rooted subnetworks identifies the full circular order of a rooted blob. We then show that subnetwork circular information is identifiable from certain data types and evolutionary models. This provides a general positive result for high-level networks, on the identifiability of the ordering in which taxon blocks attach to blobs in outer-labeled planar networks. Finally, we give examples of blobs with different internal structures which cannot be distinguished under many models and data types.
Mallows Product Measure
Published in Electron. J. Probab. 29: 1-33 (2024)
• View Publication
• BIB
Q-exchangeable ergodic distributions on the infinite symmetric group were classified by Gnedin-Olshanski (2012). In this paper, we study a specific linear combination of the ergodic measures and call it the Mallows product measure. From a particle system perspective, the Mallows product measure is a reversible stationary blocking measure of the infinite-species ASEP and it is a natural multi-species extension of the Bernoulli product blocking measures of the one-species ASEP. Moreover, the Mallows product measure can be viewed as the universal product blocking measure of interacting particle systems coming from random walks on Hecke algebras.
For the random infinite permutation distributed according to the Mallows product measure we have computed the joint distribution of its neighboring displacements, as well as several other observables. The key feature of the obtained formulas is their remarkably simple product structure. We project these formulas to ASEP with finitely many species, which in particular recovers a recent result of Adams-Balazs-Jay, and also to ASEP(q,M).
Our main tools are results of Gnedin-Olshanski about ergodic Mallows measures and shift-invariance symmetries of the stochastic colored six vertex model discovered by Borodin-Gorin-Wheeler and Galashin.
Musical Systems with $\mathbb{Z}_n$ -- Cayley Graphs
Published
• View Publication
• BIB
We apply geometric group theory to study and interpret known concepts from Western music. We show that chords, the circle of fifths, scales and certain aspects of the first species of counterpoint are encoded in the Cayley graph of the group $\mathbb{Z}_{12}$, generated by $3$ and $4$. Using $\mathbb{Z}_{12}$ as a model, we extend the above music concepts to a particular class of groups $\mathbb{Z}_{n}$, which displays geometric and algebraic features similar to $\mathbb{Z}_{12}$. We identify a weaker form of counterpoint which, in particular leads to Fux's dichotomy in $\mathbb{Z}_{12}$, and to consonant sets in $\mathbb{Z}_n$. Using Maple software, we implement these new constructions and show how to experiment with them musically.
Operad Structure of Poset Matrices
This paper examines operad structures derived from poset matrices by formulating a set of new construction rules for poset matrices. In this direction, eleven different partial composition operations will be introduced as the basis for the construction of poset matrices of any given size by extending the combinatorial setting of species of structures to poset matrices. Three of these partial composition operations are shown to define an operad structure for poset matrices. The structural properties of poset matrices and their duals are then studied based on their associated operad constructions.
The doubly asymmetric simple exclusion process, the colored Boolean process, and the restricted random growth model
The multispecies asymmetric simple exclusion process (mASEP) is a Markov chain in which particles of different species hop along a one-dimensional lattice. This paper studies the doubly asymmetric simple exclusion process $\mathrm{DASEP}(n,p,q)$ in which $q$ particles with species $1, \dots, p$ hop along a circular lattice with $n$ sites, but also the particles are allowed to spontaneously change from one species to another. In this paper, we introduce two related Markov chains called the colored Boolean process and the restricted random growth model, and we show that the DASEP lumps to the colored Boolean process, and the colored Boolean process lumps to the restricted random growth model. This allows us to generalize a theorem of David Ash on the relations between sums of steady state probabilities. We also give explicit formulas for the stationary distribution of $\mathrm{DASEP}(n,2,2)$.
Predicting Horizontal Gene Transfers with Perfect Transfer Networks
Horizontal gene transfer inference approaches are usually based on gene sequences: parametric methods search for patterns that deviate from a particular genomic signature, while phylogenetic methods use sequences to reconstruct the gene and species trees. However, it is well-known that sequences have difficulty identifying ancient transfers since mutations have enough time to erase all evidence of such events. In this work, we ask whether character-based methods can predict gene transfers. Their advantage over sequences is that homologous genes can have low DNA similarity, but still have retained enough important common motifs that allow them to have common character traits, for instance the same functional or expression profile. A phylogeny that has two separate clades that acquired the same character independently might indicate the presence of a transfer even in the absence of sequence similarity. We introduce perfect transfer networks, which are phylogenetic networks that can explain the character diversity of a set of taxa under the assumption that characters have unique births, and that once a character is gained it is rarely lost. Examples of such traits include transposable elements, biochemical markers and emergence of organelles, just to name a few. We study the differences between our model and two similar models: perfect phylogenetic networks and ancestral recombination networks. Our goals are to initiate a study on the structural and algorithmic properties of perfect transfer networks. We then show that in polynomial time, one can decide whether a given network is a valid explanation for a set of taxa, and show how, for a given tree, one can add transfer edges to it so that it explains a set of taxa. We finally provide lower and upper bounds on the number of transfers required to explain a set of taxa, in the worst case.
Phylogenetic trees defined by at most three characters
Published
• View Publication
• BIB
In evolutionary biology, phylogenetic trees are commonly inferred from a set of characters (partitions) of a collection of biological entities (e.g., species or individuals in a population). Such characters naturally arise from molecular sequences or morphological data. Interestingly, it has been known for some time that any binary phylogenetic tree can be (convexly) defined by a set of at most four characters, and that there are binary phylogenetic trees for which three characters are not enough. Thus, it is of interest to characterise those phylogenetic trees that are defined by a set of at most three characters. In this paper, we provide such a characterisation, in particular proving that a binary phylogenetic tree $T$ is defined by a set of at most three characters precisely if $T$ has no internal subtree isomorphic to a certain tree.
On the free commutative monoid over a positive operad
We study algebraic structures on the free commutative twisted algebra generated by a positive operad $\mathbf q$, in the framework of vector species. Given a nonunital commutative twisted algebra structure $μ$ on $\mathbf q$, we introduce the notion of $μ$-compatible operad structure, leading to a nonunital operad structure on $\mathbf E \circ \mathbf q$, where $\mathbf E$ stands for the exponential species. Next, we define nested pre-Lie operads (NPL-operads), a weak form of the notion of operad, in which the nested associativity axiom is weakened down to a nested pre-Lie condition. This structure is new up to our knowledge. Several constructions of NPL-operads are presented. Finally, we define algebras over a NPL-operad, based on the notion of polynomial functions.
Inferring Long-term Dynamics of Ecological Communities Using Combinatorics
In an increasingly changing world, predicting the fate of species across the globe has become a major concern. Understanding how the population dynamics of various species and communities will unfold requires predictive tools that experimental data alone can not capture. Here, we introduce our combinatorial framework, Widespread Ecological Networks and their Dynamical Signatures (WENDyS) which, using data on the relative strengths of interactions and growth rates within a community of species predicts all possible long-term outcomes of the community. To this end, WENDyS partitions the multidimensional parameter space (formed by the strengths of interactions and growth rates) into a finite number of regions, each corresponding to a unique set of coarse population dynamics. Thus, WENDyS ultimately creates a library of all possible outcomes for the community. On the one hand, our framework avoids the typical ``parameter sweeps'' that have become ubiquitous across other forms of mathematical modeling, which can be computationally expensive for ecologically realistic models and examples. On the other hand, WENDyS opens the opportunity for interdisciplinary teams to use standard experimental data (i.e., strengths of interactions and growth rates) to filter down the possible end states of a community. To demonstrate the latter, here we present a case study from the Indonesian Coral Reef. We analyze how different interactions between anemone and anemonefish species lead to alternative stable states for the coral reef community, and how competition can increase the chance of exclusion for one or more species. WENDyS, thus, can be used to anticipate ecological outcomes and test the effectiveness of management (e.g., conservation) strategies.