arXiv++ Combinatorics

Browse math.CO papers from arXiv

uniform distribution

178 papers tagged with this keyword
2019-05-23 v5
Testing Graphs against an Unknown Distribution
The area of graph property testing seeks to understand the relation between the global properties of a graph and its local statistics. In the classical model, the local statistics of a graph is defined relative to a uniform distribution over the graph's vertex set. A graph property $\mathcal{P}$ is said to be testable if the local statistics of a graph can allow one to distinguish between graphs satisfying $\mathcal{P}$ and those that are far from satisfying it. Goldreich recently introduced a generalization of this model in which one endows the vertex set of the input graph with an arbitrary and unknown distribution, and asked which of the properties that can be tested in the classical model can also be tested in this more general setting. We completely resolve this problem by giving a (surprisingly "clean") characterization of these properties. To this end, we prove a removal lemma for vertex weighted graphs which is of independent interest.
2019-05-13
Counting and sampling gene family evolutionary histories in the duplication-loss and duplication-loss-transfer models
Given a set of species whose evolution is represented by a species tree, a gene family is a group of genes having evolved from a single ancestral gene. A gene family evolves along the branches of a species tree through various mechanisms, including - but not limited to - speciation, gene duplication, gene loss, horizontal gene transfer. The reconstruction of a gene tree representing the evolution of a gene family constrained by a species tree is an important problem in phylogenomics. However, unlike in the multispecies coalescent evolutionary model, very little is known about the search space for gene family histories accounting for gene duplication, gene loss and horizontal gene transfer (the DLT-model). We introduce the notion of evolutionary histories defined as a binary ordered rooted tree describing the evolution of a gene family, constrained by a species tree in the DLT-model. We provide formal grammars describing the set of all evolutionary histories that are compatible with a given species tree, whether it is ranked or unranked. These grammars allow us, using either analytic combinatorics or dynamic programming, to efficiently compute the number of histories of a given size, and also to generate random histories of a given size under the uniform distribution. We apply these tools to obtain exact asymptotics for the number of gene family histories for two species trees, the rooted caterpillar and the complete binary tree, as well as estimates of the range of the exponential growth factor of the number of histories for random species trees of size up to 25. Our results show that including horizontal gene transfer induce a dramatic increase of the number of evolutionary histories. We also show that, within ranked species trees, the number of evolutionary histories in the DLT-model is almost independent of the species tree topology.
An Optimal Algorithm for Stopping on the Element Closest to the Center of an Interval
Real numbers from the interval [0, 1] are randomly selected with uniform distribution. There are $n$ of them and they are revealed one by one. However, we do not know their values but only their relative ranks. We want to stop on recently revealed number maximizing the probability that that number is closest to $\frac{1}{2}$. We design an optimal stopping algorithm achieving our goal and prove that its probability of success is asymptotically equivalent to $\frac{1}{\sqrt{n}}\sqrt{\frac{2}π}$.
2019-04-25 v3
The strong circular law: a combinatorial view
Let $N_n$ be an $n\times n$ complex random matrix, each of whose entries is an independent copy of a centered complex random variable $z$ with finite non-zero variance $σ^{2}$. The strong circular law, proved by Tao and Vu, states that almost surely, as $n\to \infty$, the empirical spectral distribution of $N_n/(σ\sqrt{n})$ converges to the uniform distribution on the unit disc in $\mathbb{C}$. A crucial ingredient in the proof of Tao and Vu, which uses deep ideas from additive combinatorics, is controlling the lower tail of the least singular value of the random matrix $xI - N_{n}/(σ\sqrt{n})$ (where $x\in \mathbb{C}$ is fixed) with failure probability that is inverse polynomial. In this paper, using a simple and novel approach (in particular, not using tools from additive combinatorics or any net arguments), we show that for any fixed matrix $M$ with operator norm at most $n^{0.51}$ and for all $η\geq 0$, $$\Pr\left(s_n(M+N_n) \leq η\right) \lesssim n^{C}η+ \exp(-n^{c}),$$ where $s_n(M+N_n)$ is the least singular value of $M+N_n$ and $C,c$ are absolute constants. Our result is optimal up to the constants $C,c$ and the inverse exponential-type error rate improves upon the inverse polynomial error rate due to Tao and Vu. During the course of our proof, we extend the solution of the counting problem in inverse Littlewood-Offord theory, recently isolated by the author along with Ferber, Luh, and Samotij, from Rademacher variables to general complex random variables. This significantly improves on estimates for this problem obtained using the optimal inverse Littlewood-Offord theorem of Nguyen and Vu, and may be of independent interest.
2019-04-16 v2
Extractors for small zero-fixing sources
A random variable $X$ is an $(n,k)$-zero-fixing source if for some subset $V\subseteq[n]$, $X$ is the uniform distribution on the strings $\{0,1\}^n$ that are zero on every coordinate outside of $V$. An $ε$-extractor for $(n,k)$-zero-fixing sources is a mapping $F:\{0,1\}^n\to\{0,1\}^m$, for some $m$, such that $F(X)$ is $ε$-close in statistical distance to the uniform distribution on $\{0,1\}^m$ for every $(n,k)$-zero-fixing source $X$. Zero-fixing sources were introduced by Cohen and Shinkar in [10] in connection with the previously studied extractors for bit-fixing sources. They constructed, for every $μ>0$, an efficiently computable extractor that extracts a positive fraction of entropy, i.e., $Ω(k)$ bits, from $(n,k)$-zero-fixing sources where $k\geq(\log\log n)^{2+μ}$. In this paper we present two different constructions of extractors for zero-fixing sources that are able to extract a positive fraction of entropy for $k$ essentially smaller than $\log\log n$. The first extractor works for $k\geq C\log\log\log n$, for some constant $C$. The second extractor extracts a positive fraction of entropy for $k\geq \log^{(i)}n$ for any fixed $i\in \mathbb{N}$, where $\log^{(i)}$ denotes $i$-times iterated logarithm. The fraction of extracted entropy decreases with $i$. The first extractor is a function computable in polynomial time in~$n$ (for $ε=o(1)$, but not too small); the second one is computable in polynomial time when $k\leqα\log\log n/\log\log\log n$, where $α$ is a positive constant. The subject studied in this paper is closely related to Ramsey theory. We use methods developed in Ramsey theory and our results can also be interpreted as a contribution to this field.
2019-04-15 v3
Making multigraphs simple by a sequence of double edge swaps
We show that any loopy multigraph with a graphical degree sequence can be transformed into a simple graph by a finite sequence of double edge swaps with each swap involving at least one loop or multiple edge. Our result answers a question of Janson motivated by random graph theory, and it adds to the rich literature on reachability of double edge swaps with applications in Markov chain Monte Carlo sampling from the uniform distribution of graphs with prescribed degrees.
2019-03-14
Keyed hash function from large girth expander graphs
In this paper we present an algorithm to compute keyed hash function (message authentication code MAC). Our approach uses a family of expander graphs of large girth denoted $D(n,q)$, where $n$ is a natural number bigger than one and $q$ is a prime power. Expander graphs are known to have excellent expansion properties and thus they also have very good mixing properties. All requirements for a good MAC are satisfied in our method and a discussion about collisions and preimage resistance is also part of this work. The outputs closely approximate the uniform distribution and the results we get are indistinguishable from random sequences of bits. Exact formulas for timing are given in term of number of operations per bit of input. Based on the tests, our method for implementing DMAC shows good efficiency in comparison to other techniques. 4 operations per bit of input can be achieved. The algorithm is very flexible and it works with messages of any length. Many existing algorithms output a fixed length tag, while our constructions allow generation of an arbitrary length output, which is a big advantage.
2019-02-26 v3
Polynomial bound for the partition rank vs the analytic rank of tensors
Published in Discrete Analysis, 2020:7, 18pp • View Publication • BIB
A tensor defined over a finite field $\mathbb{F}$ has low analytic rank if the distribution of its values differs significantly from the uniform distribution. An order $d$ tensor has partition rank 1 if it can be written as a product of two tensors of order less than $d$, and it has partition rank at most $k$ if it can be written as a sum of $k$ tensors of partition rank 1. In this paper, we prove that if the analytic rank of an order $d$ tensor is at most $r$, then its partition rank is at most $f(r,d,|\mathbb{F}|)$, where, for fixed $d$ and $\mathbb{F}$, $f$ is a polynomial in $r$. This is an improvement of a recent result of the author, where he obtained a tower-type bound. Prior to our work, the best known bound was an Ackermann-type function in $r$ and $d$, though it did not depend on $\mathbb{F}$. It follows from our results that a biased polynomial has low rank; there too we obtain a polynomial dependence improving the previously known Ackermann-type bound. A similar polynomial bound for the partition rank was obtained independently and simultaneously by Milićević.
2019-01-28 v2
Random graphs with given vertex degrees and switchings
Random graphs with a given degree sequence are often constructed using the configuration model, which yields a random multigraph. We may adjust this multigraph by a sequence of switchings, eventually yielding a simple graph. We show that, assuming essentially a bounded second moment of the degree distribution, this construction with the simplest types of switchings yields a simple random graph with an almost uniform distribution, in the sense that the total variation distance is $o(1)$. This construction can be used to transfer results on distributional convergence from the configuration model multigraph to the uniform random simple graph with the given vertex degrees. As examples, we give a few applications to asymptotic normality. We show also a weaker result yielding contiguity when the maximum degree is too large for the main theorem to hold.
2018-12-31 v4
Cohen-Lenstra distributions via random matrices over complete discrete valuation rings with finite residue fields
Let $(R, \mathfrak{m})$ be a complete discrete valuation ring with the finite residue field $R/\mathfrak{m} = \mathbb{F}_{q}$. Given a monic polynomial $P(t) \in R[t]$ whose reduction modulo $\mathfrak{m}$ gives an irreducible polynomial $\bar{P}(t) \in \mathbb{F}_{q}[t]$, we initiate the investigation of the distribution of $\mathrm{coker}(P(A))$, where $A \in \mathrm{Mat}_{n}(R)$ is randomly chosen with respect to the Haar probability measure on the additive group $\mathrm{Mat}_{n}(R)$ of $n \times n$ $R$-matrices. One of our main results generalizes two results of Friedman and Washington. Our other results are related to the distribution of the $\bar{P}$-part of a random matrix $\bar{A} \in \mathrm{Mat}_{n}(\mathbb{F}_{q})$ with respect to the uniform distribution, and one of them generalizes a result of Fulman. We heuristically relate our results to a celebrated conjecture of Cohen and Lenstra, which predicts that given an odd prime $p$, any finite abelian $p$-group (i.e., $\mathbb{Z}_{p}$-module) $H$ occurs as the $p$-part of the class group of a random imaginary quadratic field extension of $\mathbb{Q}$ with a probability inversely proportional to $|\mathrm{Aut}_{\mathbb{Z}}(H)|$. We review three different heuristics for the conjecture of Cohen and Lenstra, and they are all related to special cases of our main conjecture, which we prove as our main theorems. For proofs, we use some concrete combinatorial connections between $\mathrm{Mat}_{n}(R)$ and $\mathrm{Mat}_{n}(\mathbb{F}_{q})$ to translate our problems about a Haar-random matrix in $\mathrm{Mat}_{n}(R)$ into problems about a random matrix in $\mathrm{Mat}_{n}(\mathbb{F}_{q})$ with respect to the uniform distribution.
2018-11-06
The entropy of lies: playing twenty questions with a liar
`Twenty questions' is a guessing game played by two players: Bob thinks of an integer between $1$ and $n$, and Alice's goal is to recover it using a minimal number of Yes/No questions. Shannon's entropy has a natural interpretation in this context. It characterizes the average number of questions used by an optimal strategy in the distributional variant of the game: let $μ$ be a distribution over $[n]$, then the average number of questions used by an optimal strategy that recovers $x\sim μ$ is between $H(μ)$ and $H(μ)+1$. We consider an extension of this game where at most $k$ questions can be answered falsely. We extend the classical result by showing that an optimal strategy uses roughly $H(μ) + k H_2(μ)$ questions, where $H_2(μ) = \sum_x μ(x)\log\log\frac{1}{μ(x)}$. This also generalizes a result by Rivest et al. for the uniform distribution. Moreover, we design near optimal strategies that only use comparison queries of the form `$x \leq c$?' for $c\in[n]$. The usage of comparison queries lends itself naturally to the context of sorting, where we derive sorting algorithms in the presence of adversarial noise.
2018-10-12
Uniform random posets
We propose a simple algorithm generating labelled posets of given size according to the almost uniform distribution. By "almost uniform" we understand that the distribution of generated posets converges in total variation to the uniform distribution. Our method is based on a Markov chain generating directed acyclic graphs.
2018-10-04 v2
The Four Point Permutation Test for Latent Block Structure in Incidence Matrices
Transactional data may be represented as a bipartite graph $G:=(L \cup R, E)$, where $L$ denotes agents, $R$ denotes objects visible to many agents, and an edge in $E$ denotes an interaction between an agent and an object. Unsupervised learning seeks to detect block structures in the adjacency matrix $Z$ between $L$ and $R$, thus grouping together sets of agents with similar object interactions. New results on quasirandom permutations suggest a non-parametric \textbf{four point test} to measure the amount of block structure in $G$, with respect to vertex orderings on $L$ and $R$. Take disjoint 4-edge random samples, order these four edges by left endpoint, and count the relative frequencies of the $4!$ possible orderings of the right endpoint. When these orderings are equiprobable, the edge set $E$ corresponds to a quasirandom permutation $π$ of $|E|$ symbols. Total variation distance of the relative frequency vector away from the uniform distribution on 24 permutations measures the amount of block structure. Such a test statistic, based on $\lfloor |E|/4 \rfloor$ samples, is computable in $O(|E|/p)$ time on $p$ processors. Possibly block structure may be enhanced by precomputing \textbf{natural orders} on $L$ and $R$, related to the second eigenvector of graph Laplacians. In practice this takes $O(d |E|)$ time, where $d$ is the graph diameter. Five open problems are described.
2018-09-28
Low analytic rank implies low partition rank for tensors
A tensor defined over a finite field $\mathbb{F}$ has low analytic rank if the distribution of its values differs significantly from the uniform distribution. An order $d$ tensor has partition rank 1 if it can be written as a product of two tensors of order less than $d$, and it has partition rank at most $k$ if it can be written as a sum of $k$ tensors of partition rank 1. In this paper, we prove that if the analytic rank of an order $d$ tensor is at most $r$, then its partition rank is at most $f(r,d,|\mathbb{F}|)$. Previously, this was known with $f$ being an Ackermann-type function in $r$ and $d$ but not depending on $\mathbb{F}$. The novelty of our result is that $f$ has only tower-type dependence on its parameters. It follows from our results that a biased polynomial has low rank; there too we obtain a tower-type dependence improving the previously known Ackermann-type bound.
The Best-or-Worst and the Postdoc problems with random number of candidates
In this paper we consider two variants of the Secretary problem: The Best-or-Worst and the Postdoc problems. We extend previous work by considering that the number of objects is not known and follows either a discrete Uniform distribution $\mathcal{U}[1,n]$ or a Poisson distribution $\mathcal{P}(λ)$. We show that in any case the optimal strategy is a threshold strategy, we provide the optimal cutoff values and the asymptotic probabilities of success. We also put our results in relation with closely related work.
2018-06-21
A note on log-concave random graphs
Published in Electron. J. Combin. 26 (2019), no. 3, Paper No. 3.36, 9 pp • View Publication • BIB
We establish a threshold for the connectivity of certain random graphs whose (dependent) edges are determined by the uniform distributions on generalized Orlicz balls, crucially using their negative correlation properties. We also show the existence of a unique giant component for such random graphs.
2018-04-30
A large deviation principle for the Erdős-Rényi uniform random graph
Published • View Publication • BIB
Starting with the large deviation principle (LDP) for the Erdős-Rényi binomial random graph $\mathcal{G}(n,p)$ (edge indicators are i.i.d.), due to Chatterjee and Varadhan (2011), we derive the LDP for the uniform random graph $\mathcal{G}(n,m)$ (the uniform distribution over graphs with $n$ vertices and $m$ edges), at suitable $m=m_n$. Applying the latter LDP we find that tail decays for subgraph counts in $\mathcal{G}(n,m_n)$ are controlled by variational problems, which up to a constant shift, coincide with those studied by Kenyon et al. and Radin et al. in the context of constrained random graphs, e.g., the edge/triangle model.
2018-03-20 v3
A probabilistic variant of Sperner's theorem and of maximal $r$-cover free families
Published in Discrete Mathematics, October 2020, volume 343, issue 10, article 112027 • View Publication • BIB
A family of sets is called $r$-\emph{cover free} if no set in the family is contained in the union of $r$ (or less) other sets in the family. A $1$-cover free family is simply an antichain with respect to set inclusion. Thus, Sperner's classical result determines the maximal cardinality of a $1$-cover free family of subsets of an $n$-element set. Estimating the maximal cardinality of an $r$-cover free family of subsets of an $n$-element set for $r>1$ was also studied. In this note we are interested in the following probabilistic variant of this problem. Let $S_0,S_1,\ldots, S_r$ be independent and identically distributed random subsets of an $n$-element set. Which distribution minimizes the probability that $S_0\subseteq {\bigcup_{i=1}^r S_i}$? A natural candidate is the uniform distribution on an $r$-cover-free family of maximal cardinality. We show that for $r=1$ such distribution is indeed best possible. In a complete contrast, we also show that this is far from being true for every $r>1$ and $n$ large enough.
2018-02-19 v3
Further results on random cubic planar graphs
Published • View Publication • BIB
We provide precise asymptotic estimates for the number of several classes of labelled cubic planar graphs, and we analyze properties of such random graphs under the uniform distribution. This model was first analyzed by Bodirsky et al. (Random Structures Algorithms 2007). We revisit their work and obtain new results on the enumeration of cubic planar graphs and on random cubic planar graphs. In particular, we determine the exact probability of a random cubic planar graph being connected, and we show that the distribution of the number of triangles in random cubic planar graphs is asymptotically normal with linear expectation and variance. To the best of our knowledge, this is the first time one is able to determine the asymptotic distribution for the number of copies of a fixed graph containing a cycle in classes of random planar graphs arising from planar maps.
2017-09-28 v3
Hypergraph expanders from Cayley graphs
Published • View Publication • BIB
We present a simple mechanism, which can be randomised, for constructing sparse $3$-uniform hypergraphs with strong expansion properties. These hypergraphs are constructed using Cayley graphs over $\mathbb{Z}_2^t$ and have vertex degree which is polylogarithmic in the number of vertices. Their expansion properties, which are derived from the underlying Cayley graphs, include analogues of vertex and edge expansion in graphs, rapid mixing of the random walk on the edges of the skeleton graph, uniform distribution of edges on large vertex subsets and the geometric overlap property.