random sequence
28 papers tagged with this keyword
Multiple consecutive runs of multi-state trials: distributions of $(k_1, k_2, \dots, k_\ell)$ patterns
Published in Journal of Computational and Applied Mathematics, Volume 403, 113846 (2022)
• View Publication
• BIB
The pattern $(k_1, k_2, \dots, k_\ell)$ is defined to have at least $k_1$ consecutive $1$'s followed by at least $k_2$ consecutive $2$'s, $\dots$, followed by at least $k_\ell$ consecutive $\ell$'s. By iteratively applying the method that was developed previously to decouple the combinatorial complexity involved in studying complicated patterns in random sequences, the distribution of pattern $(k_1, k_2, \dots, k_\ell)$ is derived for arbitrary $\ell$. Numerical examples are provided to illustrate the results.
The m-th Longest Runs of Multivariate Random Sequences
Published in Ann Inst Stat Math 69, 497-512 (2017)
• View Publication
• BIB
The distributions of the $m$-th longest runs of multivariate random sequences are considered. For random sequences made up of $k$ kinds of letters, the lengths of the runs are sorted in two ways to give two definitions of run length ordering. In one definition, the lengths of the runs are sorted separately for each letter type. In the second definition, the lengths of all the runs are sorted together. Exact formulas are developed for the distributions of the m-th longest runs for both definitions. The derivations are based on a two-step method that is applicable to various other runs-related distributions, such as joint distributions of several letter types and multiple run lengths of a single letter type.
Joint distribution of rises, falls, and number of runs in random sequences
Published in Communications in Statistics - Theory and Methods, 48(3) (2019)
• View Publication
• BIB
By using the matrix formulation of the two-step approach to the distributions of runs, a recursive relation and an explicit expression are derived for the generating function of the joint distribution of rises and falls for multivariate random sequences in terms of generating functions of individual letters, from which the generating functions of the joint distribution of rises, falls, and number of runs are obtained. An explicit formula for the joint distribution of rises and falls with arbitrary specification is also obtained.
Distributions of successions of arbitrary multisets
Published in Communications in Statistics - Theory and Methods, 51(6), 1693-1705 (2022)
• View Publication
• BIB
By using the matrix formulation of the two-step approach to distributions of patterns in random sequences, recurrence and explicit formulas for the generating functions of successions in random permutations of arbitrary multisets are derived. Explicit formulas for the mean and variance are also obtained.
Runs in Random Sequences over Ordered Sets
Published
• View Publication
• BIB
We determine the distributions of lengths of runs in random sequences of elements from a totally ordered set (total order) or partially ordered set (partial order). In particular, we produce novel formulae for the expected value, variance, and probability generating function (PGF) of such lengths in the case of an arbitrary total order. Our focus is on the case of distributions with both atoms and diffuse (absolutely or singularly continuous) mass which has not been addressed in this generality before. We also provide a method of calculating the PGF of run lengths for countably series-parallel partial orders. Additionally, we prove a strong law of large numbers for the distribution of run lengths in a particular realization of an infinite sequence.
On the intersecting family process
Published
• View Publication
• BIB
We study the intersecting family process initially studied in \cite{BCFMR}. Here $k=k(n)$ and $E_1,E_2,\ldots,E_m$ is a random sequence of $k$-sets from $\binom{[n]}{k}$ where $E_{r+1}$ is uniformly chosen from those $k$-sets that are not already chosen and that meet $E_i,i=1,2,\ldots,r$. We prove some new results for the case where $k=cn^{1/3}$ and for the case where $k\gg n^{1/2}$.
Secretary problem and two almost the same consecutive applicants
Published
• View Publication
• BIB
We present a new variant of the secretary problem. Let $A$ be a totally ordered set of $n$ \emph{applicants}.
Given $P\subseteq A$ and $x\in A$, let $rr(P,x)=\vert\{z\in P \mid z\leq x\}\vert\mbox{ }$ be the \emph{relative rank of} $x$ \emph{with regard to} $P$, and let $rr_n(x)=rr(A,x)$. Let $x_1,x_2,\dots,x_n\in A$ be a random sequence of distinct applicants. The aim is to select $1<j\leq n$ such that $rr_n(x_{j-1})-rr_n(x_j)\in\{-1,1\}$.
Let $α$ be a real constant with $0<α<1$. Suppose the following stopping rule $τ_n(α)$: reject first $αn$ applicants and then select the first $x_j$ such that $rr(P_j,x_{j-1})-rr(P_j,x_j)\in\{-1,1\}$, where $P_j=\{x_i\mid 1\leq i\leq j\}$. Let $p_{n,τ}(α)$ be the probability that $τ_n(α)$ selects $x_j$ such that $rr_n(x_{j-1})-rr_n(x_j)\in\{-1,1\}$. We show that \[\lim_{n\rightarrow\infty}p_{n,τ}(α)\leq \lim_{n\rightarrow\infty}p_{n,τ}\left(\frac{1}{2}\right)=\frac{1}{2}\mbox{.}\]
Outliers in spectrum of sparse Wigner matrices
In this paper, we study the effect of sparsity on the appearance of outliers in the semi-circular law. Let $(W_n)_{n=1}^\infty$ be a sequence of random symmetric matrices such that each $W_n$ is $n\times n$ with i.i.d entries above and on the main diagonal equidistributed with the product $b_nξ$, where $ξ$ is a real centered uniformly bounded random variable of unit variance and $b_n$ is an independent Bernoulli random variable with a probability of success $p_n$. Assuming that $\lim\limits_{n\to\infty}n p_n=\infty$, we show that for the random sequence $(ρ_n)_{n=1}^\infty$ given by $$ρ_n:=θ_n+\frac{n p_n}{θ_n},\quad θ_n:=\sqrt{\max\big(\max\limits_{i\leq n}\|{\rm Row_i}(W_n)\|_2^2-np_n,n p_n\big)},$$ the ratio $\frac{\|W_n\|}{ρ_n}$ converges to one in probability. A non-centered counterpart of the theorem allows to obtain asymptotic expressions for eigenvalues of the Erdős--Renyi graphs, which were unknown in the regime $n p_n=Θ(\log n)$. In particular, denoting by $A_n$ the adjacency matrix of $\mathcal{G}(n,p_n)$ and by $λ_{|k|}(A_n)$ its $k$-th largest (by the absolute value) eigenvalue, under the assumptions $\lim\limits_{n\to\infty }n p_n=\infty$ and $\lim\limits_{n\to\infty}p_n=0$ we have:
-(No non-trivial outliers) If $\liminf\frac{n p_n}{\log n}\geq\frac{1}{\log (4/e)}$ then for any fixed $k\geq2$, $\frac{|λ_{|k|}(A_n)|}{2\sqrt{n p_n}}$ converges to $1$ in probability.
-(Outliers) If $\limsup\frac{n p_n}{\log n}<\frac{1}{\log (4/e)}$ then there is $\varepsilon>0$ such that for any $k\in\mathbb{N}$, we have $\lim\limits_{n\to\infty}\mathbb{P}\Big\{\frac{|λ_{|k|}(A_n)|}{2\sqrt{n p_n}}>1+\varepsilon\Big\}=1$.
On a conceptual level, our result highlights similarities in appearance of outliers in spectrum of sparse matrices and the so-called BBP phase transition phenomenon in deformed Wigner matrices.
Keyed hash function from large girth expander graphs
In this paper we present an algorithm to compute keyed hash function (message authentication code MAC). Our approach uses a family of expander graphs of large girth denoted $D(n,q)$, where $n$ is a natural number bigger than one and $q$ is a prime power. Expander graphs are known to have excellent expansion properties and thus they also have very good mixing properties. All requirements for a good MAC are satisfied in our method and a discussion about collisions and preimage resistance is also part of this work. The outputs closely approximate the uniform distribution and the results we get are indistinguishable from random sequences of bits. Exact formulas for timing are given in term of number of operations per bit of input. Based on the tests, our method for implementing DMAC shows good efficiency in comparison to other techniques. 4 operations per bit of input can be achieved. The algorithm is very flexible and it works with messages of any length. Many existing algorithms output a fixed length tag, while our constructions allow generation of an arbitrary length output, which is a big advantage.
Good weights for the Erdős discrepancy problem
Published in Discrete Analysis, 2020:8, 23 pp
• Search Publication
The Erdős discrepancy problem, now a theorem by T. Tao, asks whether every sequence with values plus or minus one has unbounded discrepancy along all homogeneous arithmetic progressions. We establish weighted variants of this problem, for weights given either by structured sequences that enjoy some irrationality features, or certain random sequences. As an intermediate result, we establish unboundedness of weighted sums of bounded multiplicative functions and products of shifts of such functions. A key ingredient in our analysis for the structured weights, is a structural result for measure preserving systems naturally associated with bounded multiplicative functions that was recently obtained in joint work with B. Host.
Biased partitions of $\mathbb{Z}^n$
Published in European Journal of Combinatorics 79 (2019): 262-270
• View Publication
• BIB
Given a function $f$ on the vertex set of some graph $G$, a scenery, let a simple random walk run over the graph and produce a sequence of values. Is it possible to, with high probability, reconstruct the scenery $f$ from this random sequence? To show this is impossible for some graphs, Gross and Grupel, call a function $f:V\to\{0,1\}$ on the vertex set of a graph $G=(V,E)$ $p$-biased if for each vertex $v$ the fraction of neighbours on which $f$ is 1 is exactly $p$. Clearly, two $p$-biased functions are indistinguishable based on their sceneries. Gross and Grupel construct $p$-biased functions on the hypercube $\{0,1\}^n$ and ask for what $p\in[0,1]$ there exist $p$-biased functions on $\mathbb{Z}^n$ and additionally how many there are. We fully answer this question by giving a complete characterization of these values of $p$. We show that $p$-biased functions exist for all $p=c/2n$ with $c\in\{0,\dots,2n\}$ and, in fact, there are uncountably many of them for every $c\in\{1,\dots,2n-1\}$. To this end, we construct uncountably many partitions of $\mathbb{Z}^n$ into $2n$ parts such that every element of $\mathbb{Z}^n$ has exactly one neighbour in each part. This additionally shows that not all sceneries on $\mathbb{Z}^n$ can be reconstructed from a sequence of values on attained on a simple random walk.
Length of the longest common subsequence between overlapping words
Published
• View Publication
• BIB
Given two random finite sequences from $[k]^n$ such that a prefix of the first sequence is a suffix of the second, we examine the length of their longest common subsequence. If $\ell$ is the length of the overlap, we prove that the expected length of an LCS is approximately $\max(\ell, \mathbb{E}[L_n])$, where $L_n$ is the length of an LCS between two independent random sequences. We also obtain tail bounds on this quantity.
A definite recursive relation and some statistical properties for Möbius function
An elementary recursive relation for M$\ddot{\mathrm{o}}$bius function $μ(n)$ is introduced by two simple ways. With this recursive relation, $μ(n)$ can be calculated without directly knowing the factorization of the $n$. $μ(1) \sim μ(2 \times 10^7) $ are calculated recursively one by one. Based on these $2\times 10^7$ samples, the empirical probabilities of $μ(n)$ of taking $-1$, 0, and 1 in classic statistics are calculated and compared with the theoretical probabilities in number theory. The numerical consistency between these two kinds of probability show that $μ(n)$ could be seen as an independent random sequence when $n$ is large. The expectation and variance of the $μ(n)$ are $0$ and $6 n/ π^2$, respectively. Furthermore, we show that any conjecture of the Mertens type is false in probability sense, and present an upper bound for cumulative sums of $μ(n)$ with a certain probability.
The sharp threshold for making squares
Published
• View Publication
• BIB
Consider a random sequence of $N$ integers, each chosen uniformly and independently from the set $\{1,\dots,x\}$. Motivated by applications to factorisation algorithms such as Dixon's algorithm, the quadratic sieve, and the number field sieve, Pomerance in 1994 posed the following problem: how large should $N$ be so that, with high probability, this sequence contains a subsequence, the product of whose elements is a perfect square? Pomerance determined asymptotically the logarithm of the threshold for this event, and conjectured that it in fact exhibits a sharp threshold in $N$. More recently, Croot, Granville, Pemantle and Tetali determined the threshold up to a factor of $4/π+ o(1)$ as $x \to \infty$, and made a conjecture regarding the location of the sharp threshold.
In this paper we prove both of these conjectures, by determining the sharp threshold for making squares. Our proof combines techniques from combinatorics, probability and analytic number theory; in particular, we use the so-called method of self-correcting martingales in order to control the size of the 2-core of the random hypergraph that encodes the prime factors of our random numbers. Our method also gives a new (and completely different) proof of the upper bound in the main theorem of Croot, Granville, Pemantle and Tetali.
Aperiodic Crosscorrelation of Sequences Derived from Characters
Published
• View Publication
• BIB
It is shown that pairs of maximal linear recursive sequences (m-sequences) typically have mean square aperiodic crosscorrelation on par with that of random sequences, but that if one takes a pair of m-sequences where one is the reverse of the other, and shifts them appropriately, one can get significantly lower mean square aperiodic crosscorrelation. Sequence pairs with even lower mean square aperiodic crosscorrelation are constructed by taking a Legendre sequence, cyclically shifting it, and then cutting it (approximately) in half and using the halves as the sequences of the pair. In some of these constructions, the mean square aperiodic crosscorrelation can be lowered further if one truncates or periodically extends (appends) the sequences. Exact asymptotic formulae for mean squared aperiodic crosscorrelation are proved for sequences derived from additive characters (including m-sequences and modified versions thereof) and multiplicative characters (including Legendre sequences and their relatives). Data is presented that shows that sequences of modest length have performance that closely approximates the asymptotic formulae.
Low Correlation Sequences from Linear Combinations of Characters
Published
• View Publication
• BIB
Pairs of binary sequences formed using linear combinations of multiplicative characters of finite fields are exhibited that, when compared to random sequence pairs, simultaneously achieve significantly lower mean square autocorrelation values (for each sequence in the pair) and significantly lower mean square crosscorrelation values. If we define crosscorrelation merit factor analogously to the usual merit factor for autocorrelation, and if we define demerit factor as the reciprocal of merit factor, then randomly selected binary sequence pairs are known to have an average crosscorrelation demerit factor of $1$. Our constructions provide sequence pairs with crosscorrelation demerit factor significantly less than $1$, and at the same time, the autocorrelation demerit factors of the individual sequences can also be made significantly less than $1$ (which also indicates better than average performance). The sequence pairs studied here provide combinations of autocorrelation and crosscorrelation performance that are not achievable using sequences formed from single characters, such as maximal linear recursive sequences (m-sequences) and Legendre sequences. In this study, exact asymptotic formulae are proved for the autocorrelation and crosscorrelation merit factors of sequence pairs formed using linear combinations of multiplicative characters. Data is presented that shows that the asymptotic behavior is closely approximated by sequences of modest length.
Sparse exchangeable graphs and their limits via graphon processes
Published in Journal of Machine Learning Research 18(210):1-71, 2018
• Search Publication
In a recent paper, Caron and Fox suggest a probabilistic model for sparse graphs which are exchangeable when associating each vertex with a time parameter in $\mathbb{R}_+$. Here we show that by generalizing the classical definition of graphons as functions over probability spaces to functions over $σ$-finite measure spaces, we can model a large family of exchangeable graphs, including the Caron-Fox graphs and the traditional exchangeable dense graphs as special cases. Explicitly, modelling the underlying space of features by a $σ$-finite measure space $(S,\mathcal{S},μ)$ and the connection probabilities by an integrable function $W\colon S\times S\to [0,1]$, we construct a random family $(G_t)_{t\geq 0}$ of growing graphs such that the vertices of $G_t$ are given by a Poisson point process on $S$ with intensity $tμ$, with two points $x,y$ of the point process connected with probability $W(x,y)$. We call such a random family a graphon process. We prove that a graphon process has convergent subgraph frequencies (with possibly infinite limits) and that, in the natural extension of the cut metric to our setting, the sequence converges to the generating graphon. We also show that the underlying graphon is identifiable only as an equivalence class over graphons with cut distance zero. More generally, we study metric convergence for arbitrary (not necessarily random) sequences of graphs, and show that a sequence of graphs has a convergent subsequence if and only if it has a subsequence satisfying a property we call uniform regularity of tails. Finally, we prove that every graphon is equivalent to a graphon on $\mathbb{R}_+$ equipped with Lebesgue measure.
Dimers on Rail Yard Graphs
Published in Ann. Inst. Henri Poincaré Comb. Phys. Interact. 4 (2017), 479-539
• View Publication
• BIB
We introduce a general model of dimer coverings of certain plane bipartite graphs, which we call rail yard graphs (RYG). The transfer matrices used to compute the partition function are shown to be isomorphic to certain operators arising in the so-called boson-fermion correspondence. This allows to reformulate the RYG dimer model as a Schur process, i.e. as a random sequence of integer partitions subject to some interlacing conditions.
Beyond the computation of the partition function, we provide an explicit expression for all correlation functions or, equivalently, for the inverse Kasteleyn matrix of the RYG dimer model. This expression, which is amenable to asymptotic analysis, follows from an exact combinatorial description of the operators localizing dimers in the transfer-matrix formalism, and then a suitable application of Wick's theorem.
Plane partitions, domino tilings of the Aztec diamond, pyramid partitions, and steep tilings arise as particular cases of the RYG dimer model. For the Aztec diamond, we provide new derivations of the edge-probability generating function, of the biased creation rate, of the inverse Kasteleyn matrix and of the arctic circle theorem.
Variances and Covariances in the Central Limit Theorem for the Output of a Transducer
Published in European J. Combin. 49 (2015), 167--187
• View Publication
• BIB
We study the joint distribution of the input sum and the output sum of a deterministic transducer. Here, the input of this finite-state machine is a uniformly distributed random sequence.
We give a simple combinatorial characterization of transducers for which the output sum has bounded variance, and we also provide algebraic and combinatorial characterizations of transducers for which the covariance of input and output sum is bounded, so that the two are asymptotically independent.
Our results are illustrated by several examples, such as transducers that count specific blocks in the binary expansion, the transducer that computes the Gray code, or the transducer that computes the Hamming weight of the width-$w$ non-adjacent form digit expansion. The latter two turn out to be examples of asymptotic independence.
Geometric approach to string analysis: deviation from linearity and its use for biosequence classification
Published
• View Publication
• BIB
Tools that effectively analyze and compare sequences are of great importance in various areas of applied computational research, especially in the framework of molecular biology. In the present paper, we introduce simple geometric criteria based on the notion of string linearity and use them to compare DNA sequences of various organisms, as well as to distinguish them from random sequences. Our experiments reveal a significant difference between biosequences and random sequences - the former having much higher deviation from linearity than the latter - as well as a general trend of increasing deviation from linearity between primitive and biologically complex organisms.