sequence
6966 papers tagged with this keyword
Simplifying and Characterizing DAGs and Phylogenetic Networks via Least Common Ancestor Constraints
Published
• View Publication
• BIB
Rooted phylogenetic networks, or more generally, directed acyclic graphs (DAGs), are widely used to model species or gene relationships that traditional rooted trees cannot fully capture, especially in the presence of reticulate processes or horizontal gene transfers. Such networks or DAGs are typically inferred from observable data (e.g. genomic sequences of extant species), providing only an estimate of the true evolutionary history. However, these inferred DAGs are often complex and difficult to interpret. In particular, many contain vertices that do not serve as least common ancestors (LCAs) for any subset of the underlying genes or species, thus may lack direct support from the observable data. In contrast, LCA vertices are witnessed by historical traces justifying their existence and thus represent ancestral states substantiated by the data. To reduce unnecessary complexity and eliminate unsupported vertices, we aim to simplify a DAG to retain only LCA vertices while preserving essential evolutionary information.
In this paper, we characterize $\mathrm{LCA}$-relevant and $\mathrm{lca}$-relevant DAGs, defined as those in which every vertex serves as an LCA (or unique LCA) for some subset of taxa. We introduce methods to identify LCAs in DAGs and efficiently transform any DAG into an $\mathrm{LCA}$-relevant or $\mathrm{lca}$-relevant one while preserving key structural properties of the original DAG or network. This transformation is achieved using a simple operator ``$\ominus$'' that mimics vertex suppression.
Asymptotic Bounds and Online Algorithms for Average-Case Matrix Discrepancy
We study the matrix discrepancy problem in the average-case setting. Given a sequence of $m \times m$ symmetric matrices $A_1,\ldots,A_n$, its discrepancy is defined as the minimal spectral norm over all signed sums $\sum_{i=1}^n x_iA_i$ with $x_1,\ldots,x_n \in \{\pm1\}$. Our contributions are twofold. First, we study the asymptotic discrepancy of random matrices. When the matrices belong to the Gaussian orthogonal ensemble, we provide a sharp characterization of the asymptotic discrepancy and show that the limiting distribution is concentrated around $Θ(\sqrt{nm}4^{-(1 + o(1))n/m^2})$, under the assumption $m^2 \ll n/\log{n}$. We observe that the trivial bound $O(\sqrt{nm})$ cannot be improved when $n \ll m^2$ and show that this phenomenon occurs for a broad class of random matrices. In the case $n = Ω(m^2)$, we provide a matching upper bound. Second, we analyse the matrix hyperbolic cosine algorithm, an online algorithm for matrix discrepancy minimization due to Zouzias (2011), in the average-case setting. We show that the algorithm achieves with high probability a discrepancy of $O(m\log{m})$ for a broad class of random matrices, including Wigner matrices with entries satisfying a hypercontractive inequality and Gaussian Wishart matrices.
Periodic orbits on 2-regular circulant digraphs
Published
• View Publication
• BIB
Periodic orbits (equivalence classes of closed paths up to cyclic shifts) play an important role in applications of graph theory. For example, they appear in the definition of the Ihara zeta function and exact trace formulae for the spectra of quantum graphs. Circulant graphs are Cayley graphs of $\mathbb{Z}_n$. Here we consider directed Cayley graphs with two generators (2-regular Cayley digraphs). We determine the number of primitive periodic orbits of a given length (total number of directed edges) in terms of the number of times edges corresponding to each generator appear in the periodic orbit (the step count). Primitive periodic orbits are those periodic orbits that cannot be written as a repetition of a shorter orbit. We describe the lattice structure of lengths and step counts for which periodic orbits exist and characterize the repetition number of a periodic orbit by its winding number (the sum of the step sequence divided by the number of vertices) and the repetition number of its step sequence. To obtain these results, we also evaluate the number of Lyndon words on an alphabet of two letters with a given length and letter count.
Combinatorial connections in snake graphs: Tilings, lattice paths, and perfect matchings
Snake graphs and their perfect matchings play a key role in the description of cluster variables of cluster algebras associated to surfaces. In this paper, we introduce triangular snake graphs and establish a bijection between their routes (non-intersecting lattice paths), perfect matchings of their underlying snake graphs, and tilings. As an application, we show that the number of perfect matchings in straight snake graphs can be expressed in terms of determinants of Hankel matrices with Catalan number entries. Moreover, we prove that the number of perfect matchings in snake graphs can be expressed as a sum of products of Fibonacci numbers, and we show how Fibonacci and Pell sequences arise from determinants of matrices with Fibonacci entries.
Dual Mixed Volume
Published
• View Publication
• BIB
We define and study the dual mixed volume rational function of a sequence of polytopes, a dual version of the mixed volume polynomial. This concept has direct relations to the adjoint polynomials and the canonical forms of polytopes. We show that dual mixed volume is additive under mixed subdivisions, and is related by a change of variables to the dual volume of the Cayley polytope. We study dual mixed volume of zonotopes, generalized permutohedra, and associahedra. The latter reproduces the planar $φ^3$-scalar amplitude at tree level.
Berge Pancyclic hypergraphs
Published
• View Publication
• BIB
A Berge cycle of length $\ell$ in a hypergraph is an alternating sequence of $\ell$ distinct vertices and $\ell$ distinct edges $v_1,e_1,v_2, \ldots, v_\ell, e_{\ell}$ such that $\{v_i, v_{i+1}\} \subseteq e_i$ for all $i$, with indices taken modulo $\ell$. We call an $n$-vertex hypergraph pancyclic if it contains Berge cycles of every length from $3$ to $n$. We prove a sharp Dirac-type result guaranteeing pancyclicity in uniform hypergraphs: for $n \geq 70$, $3 \leq r \leq \lfloor (n-1)/2\rfloor - 2$, if $\cH$ is an $n$-vertex, $r$-uniform hypergraph with minimum degree at least ${\lfloor (n-1)/2 \rfloor \choose r-1} + 1$, then $\cH$ is pancyclic.
Adjacent cycle-chains are $e$-positive
Published
• View Publication
• BIB
We describe a way to decompose the chromatic symmetric function as a positive sum of smaller pieces. We show that these pieces are $e$-positive for cycles. Then we prove that attaching a cycle to a graph preserves the $e$-positivity of these pieces. From this, we prove an $e$-positive formula for graphs of cycles connected at adjacent vertices. We extend these results to graphs formed by connecting a sequence of cycles and cliques.
Bijections for generalized Wilf equivalences
Published
• View Publication
• BIB
Starting with an inclusion-exclusion proof of a combinatorial identity, a direct bijection can be produced using recursive subtraction (sometimes with a direct combinatorial description). We apply this method to identities for generalized Wilf equivalences among consecutive patterns in inversion sequences, giving direct bijective proofs of some generalized Wilf equivalences shown by Auli and Elizalde. We also give new bijective proofs of a stronger relation among some consecutive patterns.
Multifold Convolutions, Generating Functions and 1d Random Walks
We consider multifold convolutions of a combinatorial sequence $(a_n)_{n=0}^{\infty}$: namely, for each $k \in \N$ the $k$-fold convolution is $\mathcal{M}^{(k)}_n(\boldsymbol{a}) = \sum_{j_1+\dots+j_k=n} a_{j_1} \cdots a_{j_k}$. Let $C_n$ be the Catalan numbers, and let $B_n$ be the central binomial coefficients. Then for random Dyck paths or simple random walk bridges, the multifold convolutions give moments of returns to the origin, using the stars-and-bars problem. There are well-known explicit formulas for the multifold convolutions of $C_n$ and $B_n$. But even for combinatorial sequences $B_n^2$ and $B_n^3$, one may determine asymptotics of multifold convolutions for large $n$. We also discuss large deviations: In a second part of the paper we consider an elementary version of the circle method for calculating asymptotics using complex analysis.
Cryptarithmically unique terms in integer sequences
A cryptarithm (or alphametic) is a mathematical puzzle in which numbers are represented with words in such a way that identical letters stand for equal digits and distinct letters for unequal digits. An alphametic puzzle is usually given in the form of an equation that needs to be solved, such as SEND + MORE = MONEY. Alternatively, here we will consider cryptarithms constrained not by an equation but by a particular subsequence of natural numbers, for example perfect squares or primes. Such a cryptarithm has a unique solution if there is exactly one term in the sequence that has the corresponding pattern of digits. We will call such terms cryptarithmically unique. Here we estimate the density of such terms in an arbitrary sequence for which the overall density of terms among integers is known. In particular, among all perfect squares below 10^12, slightly less than one half are cryptarithmically unique, their density increasing toward larger numbers. Cryptarithmically unique prime numbers, however, are initially very scarce. Combinatorial estimates suggest that their density should drop below 10^-300 for decimal lengths of approximately 1829 digits, but then it recovers and is asymptotic to unity for very large primes. Finally, we introduce and discuss primonumerophobic digit patterns that no prime number happens to have.
Centralizers in the plactic monoid
Published
• View Publication
• BIB
Let u be a word over the positive integers. Motivated in part by a question from representation theory, we study the centralizer set of u which is C(u) = {w | uw is Knuth-equivalent to wu}. In particular, we give various necessary conditions for w to be in C(u). We also characterize C(u) when u has few letters, when it has a single repeated entry, or when it is a certain type of decreasing sequence. We consider c_{n,m}(u), the number of w in C(u) of length n with max w at most m. We prove that for |u| = 1 the value of this function depends only on the relative sizes of u and m and not on their actual values. And for various u we use Stanley's theory of poset partitions to show that, for fixed n, c_{n,m}(u) is a polynomial in m with certain degree and leading coefficient. We end with various conjectures and directions for further research.
Garsia--Remmel $q$-rook numbers are not always unimodal
Published
• View Publication
• BIB
We show by an explicit example that the Garsia--Remmel $q$-rook numbers of Ferrers boards do not all have unimodal sequences of coefficients. This resolves in the negative a question from 1986 by the aforementioned authors.
Counting independent sets in regular graphs with bounded independence number
An $n$-vertex, $d$-regular graph can have at most $2^{n/2+o_d(n)}$ independent sets. In this paper we address what happens with this upper bound when we impose the further condition that the graph has independence number at most $α$.
We give upper and lower bounds that in many cases are close to each other. In particular, for each $0 < c_{\rm ind} \leq 1/2$ we exhibit a constant $k(c_{\rm ind})$ such that if $(G_n)_{n \in {\mathbb N}}$ is a sequence of graphs with $G_n$ $d$-regular on $n$ vertices and with maximum independent set size at most $α$, with $d\rightarrow \infty$ and $α/n \rightarrow c_{\rm ind}$ as $n \rightarrow \infty$, then $G_n$ has at most $k(c_{\rm ind})^{n+o(n)}$ independent sets, and we show that there is a sequence $(G_n)_{n \in {\mathbb N}}$ of graphs with $G_n$ $d$-regular on $n$ vertices ($d \leq n/2$) and with maximum independent set size at most $α$, with $α/n \rightarrow c_{\rm ind}$ as $n \rightarrow \infty$ and with $G_n$ having at least $k(c_{\rm ind})^{n+o(n)}$ independent sets. We also consider the regime $1/2 < c_{\rm ind} < 1$. Here for each $0 < c_{\rm deg} \leq 1-c_{\rm ind}$ we exhibit a constant $k(c_{\rm ind},c_{\rm deg})$ for which an analogous pair of statements can be proven, except that in each case we add the condition $d/n \rightarrow c_{\rm deg}$ as $n \rightarrow \infty$.
Our upper bounds are based on graph container arguments, while our lower bounds are constructive.
On maximal almost balanced non-overlapping codes and non-overlapping codes with restricted run-lengths
Published
• View Publication
• BIB
This paper concerns non-overlapping codes, block codes motivated by synchronisation and DNA-based storage applications. Most existing constructions of these codes do not account for the restrictions posed by the physical properties of communication channels. If undesired sequences are not avoided, the system using the encoding may start behaving incorrectly. Hence, we aim to characterise all non-overlapping codes satisfying two additional constraints. For the first constraint, where approximately half of the letters in each word are positive, we derive necessary and sufficient conditions for the code's non-expandability and improve known bounds on its maximum size. We also determine exact values for the maximum sizes of polarity-balanced non-overlapping codes having small block and alphabet sizes. For the other constraint, where long sequences of consecutive equal symbols lead to undesired behaviour, we derive bounds and constructions of constrained non-overlapping codes. Moreover, we provide constructions of non-overlapping codes that satisfy both constraints and analyse the sizes of the obtained codes.
Growth of recurrences with mixed multifold convolutions
Generalizing some popular sequences like Catalan's number, Schröder's number, etc, we consider the sequence $s_n$ with $s_0=1$ and for $n\ge 1$, \begin{multline*}
s_n=\sum_{x_1+\dots+x_{\ell_1}=n-1} κ_1 s_{x_1}\dots s_{x_{\ell_1}} + \dots +\sum_{x_1+\dots+x_{\ell_{t'}}=n-1} κ_{t'} s_{x_1}\dots s_{x_{\ell_{t'}}}+\\ \max_{x_1+\dots+x_{\ell_{t'+1}}=n-1} κ_{t'+1} s_{x_1}\dots s_{x_{\ell_{t'+1}}} + \dots + \max_{x_1+\dots+x_{\ell_t}=n-1} κ_t s_{x_1}\dots s_{x_{\ell_t}}, \end{multline*} where $x_i$ are nonnegative integers, $\ell_1,\dots,\ell_t$ are positive integers, and $κ_1,\dots,κ_t$ are positive reals. We show that it is possible to compute the growth rate $λ$ of $s_n$ to any precision. In particular, for every $n\ge 2$,
\[
\sqrt[n]{\frac{κ^*}{\mathcal L(n-1) s_1} s_n} \le λ\le \sqrt[n]{3^{18\log 3 + 2\log\frac{s_1\mathcal L^2}{κ^*}} n^{3\log n + 12\log 3 + \log\frac{s_1\mathcal L^2}{κ^*}} s_n},
\]where $\mathcal L=\max_i \ell_i$ and $κ^*=κ_i$ for some $i$ with $\ell_i\ge 2$, and the logarithm has the base $\frac{\mathcal L+1}{\mathcal L}$. The constants in the inequalities are not very well optimized and serve mostly as a proof of concept with the ratio of the upper bound and the lower bound converging to $1$ as $n$ goes to infinity.
Asymptotic Normality and Concentration Inequalities of Statistics of Core Partitions with Bounded Perimeters
Core partitions have attracted much attention since Anderson's work (2002) on the number of $(s,t)$-core partitions for coprime $s,t$. Recently, there has been a growing interest in studying the limiting distributions of the sizes of random simultaneous core partitions. In this paper, we prove the asymptotic normality of certain statistics of uniform random core partitions with bounded perimeters in the Kolmogorov and Wasserstein $W_1$ distances, including the length and size of a random (strict) $n$-core partition, the length of the Durfee square and the size of a random self-conjugate $n$-core partition. Accordingly, we prove that these statistics are subgaussian. This contrasts with the asymptotic behavior of the size of a random $(s, t)$-core partition for coprime $s,t$ studied by Even-Zohar (2022), which converges in law to Watson's $U^2$ distribution. Our results show that the distribution of the size of a random strict $(n, dn+1)$-core partition is asymptotically normal when $d \ge 3$ is fixed and $n$ tends to infinity, which is an analog of Zaleski's conjecture (2017) and covers Komlós, Sergel, and Tusnády's result (2020) as a special case. Our proof integrates a variety of combinatorial and probabilistic tools, including Stein's method based on Hoeffding decomposition, Hoeffding's combinatorial central limit theorem, the Efron-Stein inequalities on product spaces and slices, and asymptotics of Pólya frequency sequences. Furthermore, our approach is potentially applicable to the study of the asymptotic normality of functionals of random variables with certain global dependence structures that can be decomposed into appropriate mixture forms.
Limits of sparse hypergraphs
We generalize ultraproducts and local-global limits of graphs to hypergraphs and other structures. We show that the local statistics of an ultraproduct of a sequence of hypergraphs are the ultralimits of the local statistics of the hypergraphs. Using some standard results from model theory, we conclude that the space of (equivalence classes of) pmp hypergraphs with the topology of local-global convergence is compact, and that any countable set of local statistics for a pmp hypergraph can be realized as the statistics of a set of labellings (rather than just approximated) in a local-global equivalent hypergraph.
We give two applications. First, we characterize those structures where any solution to the corresponding CSP can be turned into a measurable solution. These turn out to be the width-1 structures. We can also use the limit machinery to extract from this theorem a purely finitary characterizations of width-1 structures involving asymptotic solutions.
Second, we prove two measurable versions of the Frankl--Rödl matching theorem using measurable nibble and differential equation arguments. The measurable proofs are much softer than the purely finitary results. And, we can recover the finitary theorems using the limit machinery.
Faber-Krahn type inequality for supertrees
Published
• View Publication
• BIB
The Faber-Krahn inequality states that the first Dirichlet eigenvalue among all bounded domains is no less than a Euclidean ball with the same volume in $\mathbb{R}^n$ \cite{Chavel FB}. Bıyıkoğlu and Leydold (J. Comb. Theory, Ser. B., 2007) demonstrated that the Faber-Krahn inequality also holds for the class of trees with boundary with the same degree sequence and characterized the unique extremal tree. Bıyıkoğlu and Leydold (2007) also posed a question as follows: Give a characterization of all graphs in a given class $\mathcal{C}$ with the Faber-Krahn property. In this paper, we address this question specifically for $k$-uniform supertrees with boundary. We introduce a spiral-like ordering (SLO-ordering) of vertices for supertrees, an extension of the SLO-ordering for trees initially proposed by Pruss [ Duke Math. J., 1998], and prove that the SLO-supertree has the Faber-Krahn property among all supertrees with a given degree sequence. Furthermore, among degree sequences that have a minimum degree $d$ for interior vertices, the SLO-supertree with degree sequence $(d,\ldots,d, d', 1, \dots, 1)$ possesses the Faber-Krahn property.
Average-case matrix discrepancy: satisfiability bounds
Published
• View Publication
• BIB
Given a sequence of $d \times d$ symmetric matrices $\{\mathbf{W}_i\}_{i=1}^n$, and a margin $Δ> 0$, we investigate whether it is possible to find signs $(ε_1, \dots, ε_n) \in \{\pm 1\}^n$ such that the operator norm of the signed sum satisfies $\|\sum_{i=1}^n ε_i \mathbf{W}_i\|_{\rm op} \leq Δ$. Kunisky and Zhang (2023) recently introduced a random version of this problem, where the matrices $\{\mathbf{W}_i\}_{i=1}^n$ are drawn from the Gaussian orthogonal ensemble. This model can be seen as a random variant of the celebrated Matrix Spencer conjecture and as a matrix-valued analog of the symmetric binary perceptron in statistical physics. In this work, we establish a satisfiability transition in this problem as $n, d \to \infty$ with $n / d^2 \to τ> 0$. First, we prove that the expected number of solutions with margin $Δ=κ\sqrt{n}$ has a sharp threshold at a critical $τ_1(κ)$: for $τ< τ_1(κ)$ the problem is typically unsatisfiable, while for $τ> τ_1(κ)$ the average number of solutions is exponentially large. Second, combining a second-moment method with recent results from Altschuler (2023) on margin concentration in perceptron-type problems, we identify a second threshold $τ_2(κ)$, such that for $τ>τ_2(κ)$ the problem admits solutions with high probability. In particular, we establish that a system of $n = Θ(d^2)$ Gaussian random matrices can be balanced so that the spectrum of the resulting matrix macroscopically shrinks compared to the semicircle law. Finally, under a technical assumption, we show that there exists values of $(τ,κ)$ for which the number of solutions has large variance, implying the failure of the second moment method. Our proofs rely on establishing concentration and large deviation properties of correlated Gaussian matrices under spectral norm constraints.
Caterpillars with given degree sequence, small Energy and small Hosoya index
The energy $En(G)$ of a graph $G$ is defined as the sum of the absolute values of its eigenvalues. The Hosoya index $Z(G)$ of a graph $G$ is the number of independent edge subsets of $G$, including the empty set. For any given degree sequence $D$, we characterize the caterpillar $\mathcal{S}(D)$ that has the minimum $Z$ and $En$. %and maximum $σ$. In $\mathcal{S}(D)$, as we move along the internal path towards the center, large and small degrees alternate. We also compare $\mathcal{S}(D)$ with $\mathcal{S}(Y)$, for a degree sequence $Y$ majorized by a degree sequence $D$. Suppose $Y=(y_1,\dots ,y_n)$ and $D=(d_1,\dots ,d_n)$ are degree sequences such that $Y$ is majorized by $D$ and$$\sum_{i=1}^{n}y_i=\sum_{i=1}^{n}d_i,$$then $Z(\mathcal{S}(D))<Z(\mathcal{S}(Y))$ and $En(\mathcal{S}(D))<En(\mathcal{S}(Y))$.