arXiv++ Combinatorics

Browse math.CO papers from arXiv

math.ST ↗ arXiv

25 papers in this category
2026-09-22
Concentration of Regularized Sparse Random Matrices: Spectral Edge Bounds via Nonbacktracking Operators
In sparse random matrices, spectral outliers (eigenvalues and singular values located away from the bulk) emerge due to degree fluctuations: high degrees inflate the operator norm, while low column degrees reduce the least singular value. As proved by Feige and Ofek (2005) and Le, Levina, and Vershynin (2017), degree regularization enforces concentration at the expected norm scale. However, precise bounds incorporating the cutoffs remain unexplored and challenging since regularization introduces dependencies among entries. For the first time in the literature, we provide variance- and cutoff-dependent bounds for extreme singular values and eigenvalues of regularized inhomogeneous random matrices. In the absence of regularization, our lower bound for the least singular value matches the same leading constant obtained by Brailovskaya and van Handel (2024). Moreover, our error term vanishes under the milder condition $d/\log N\to\infty$, as opposed to their stronger requirement $d/(\log N)^4\to\infty$. A key ingredient is to extend spectral radius bounds for nonbacktracking matrices to the dependent setting. We build on approaches for independent cases established by Benaych-Georges, Bordenave, and Knowles (2020), as well as Dumitriu and Zhu (2024), and carefully handle edges traversed only once. Our proof framework separates deterministic spectral comparisons from probabilistic estimates: once Loewner inequalities and columnwise variance controls are established, the remaining probabilistic analysis boils down to verifying the graph moment conditions formulated in this paper. We hope this framework can be extended to handle general random matrices with more complex dependencies.
2026-09-20 v7
New matrix perturbation bounds with relative strength: Perturbation of eigenspaces
Matrix perturbation bounds (such as Weyl and Davis--Kahan) are used abundantly in many areas of mathematics and data science. Many bounds (such as the above two) involve the spectral norm of the noise matrix and are sharp in worst-case analysis. In order to refine these classical bounds, we introduce a new parameter, which we refer to as the relative strength. This parameter measures the strength of the action of the noise matrix on the relevant eigenvectors of the ground matrix. It has turned out that in a number of situations, we can use the relative strength as a replacement for the spectral norm (which can be seen as the absolute strength). This has led to a number of notable improvements under certain sets of assumptions, which are frequently met in practice. A representative example is the case when the noise matrix is random. For the purpose of our study, we introduce a new method of analysis, which combines the classical contour integral argument with new (combinatorial) ideas. This method is robust and of independent interest. In the current paper, we focus on the perturbation of eigenspaces (Davis--Kahan type results). Perturbation bounds for eigenspaces are essential in statistics and theoretical computer science, and thus deserve a special treatment. Furthermore, this will lay the ground for the more technical treatment of general matrix functionals, which appears in a future paper.
2026-09-17
Sharp spectral norm concentration of sparse random tensors
We prove a sharp concentration inequality for the spectral norm of sparse random tensors with independent Bernoulli entries. Let $T$ be an order-$k$ tensor of dimension $n\times\cdots\times n$ with independent Bernoulli$(p)$ entries, where $k$ is fixed. For any $c,r>0$, we show that $\|T-\mathbb E T\|\le C_{k,r,c}\sqrt{np}$ with probability at least $1-n^{-r}$ whenever $np\ge c\log n$. We extend this bound to inhomogeneous Bernoulli sampling with deterministic entrywise weights. This removes the logarithmic factor in the work of Zhou and Zhu (2021). The proof follows the Kahn--Szemerédi light--heavy decomposition with a refined estimate on the heavy tuple part. We also obtain a log-free second eigenvalue bound for the random hypergraph model of Friedman and Wigderson (1995).
2026-09-15
An Alon-Boppana Bound for the Non-Backtracking Operator
For any fixed $k$, we prove a lower bound on the $k$th largest modulus of an eigenvalue of the non-backtracking matrix $B$. Specifically, consider any deterministic or random family of graphs that converges locally to the unimodular Galton-Watson tree with root degree distribution $D$, and set $κ:=\mathbb E[D(D-1)]/\mathbb E[D]$. Given $κ>1$ and an exponential-moment bound on the empirical degree distributions, we show that $|λ_k(B)|\geq\sqrtκ-o_N(1)$, where $N$ is the number of vertices. When restricted to locally tree-like regular graphs, this recovers a well-known consequence of the Ihara-Bass formula. In the specific case where the graph is generated through the Erdős-Rényi model with expected degree $d>1$, this proves a conjecture of Bordenave, Lelarge, and Massoulié. To do this, we show that the normalized log-determinant of the Bethe-Hessian of the graph is bounded by that of the Bethe-Hessian of its local limit. This bound is violated if the eigenvalues of the non-backtracking matrix are too small. We establish this using an effective-conductance interpretation of the tree Green's function recursion.
2026-09-10
Identifiability of Nonnegative Tensor Decompositions via Positive Scattering
Identifiability of tensor decompositions is often established through linear-algebraic conditions on the factor families. For nonnegative decompositions, however, positivity provides additional information that is not captured by dimension and independence alone: nonnegative terms cannot cancel, and their supports constrain competing decompositions. We introduce a positive scattering term that quantifies this additional source of identifiability and combine it with the dimension budget underlying the Lovitz--Petrov generalization of Kruskal's theorem. For every subset of components, we obtain two sufficient conditions: a threshold of $2|S|-2$ guarantees minimality and nonnegative rank, while the stronger threshold $2|S|-1$ guarantees uniqueness among nonnegative decompositions of the same length. The key result is a positive splitting inequality for irreducible exchanges of nonnegative rank-one tensors, which combines the dimension constraint with support-induced geometric rigidity. Although the scattering term is defined through an optimization over intermediate factor spaces, we show that its mode costs are exactly $0$, $1$, or $+\infty$, yielding an exact activation characterization in terms of graph connectivity. The resulting criterion can strictly certify sparse nonnegative tensor decompositions beyond the reach of Kruskal and Lovitz--Petrov conditions, including examples for which those conditions fail even after reshaping. In the matrix case, the two criteria reduce respectively to full-rank factorization and two-sided separability.
2026-09-07
Tensor network representations of discrete maximum entropy distributions via mean polytopes
We present tensor network representations for discrete maximum entropy distributions under expectation constraints. To this end, we introduce Computation-Activation Networks (CompActNets), a tensor network architecture that subsumes exponential families. By leveraging the geometry of the convex polytope of realizable expectation vectors, we represent any maximum entropy distribution in the same architecture. We exploit the fact that proper faces of this polytope correspond to the boundary closure of exponential families, which restricts the distribution's support. We then derive explicit representations for the support within the CompActNet architecture. The proposed framework suggests tensor network ranks as complexity measures for faces. Finally, a case study on Boolean statistics links the geometry of 0/1-polytopes directly to propositional formulas.
2026-08-17
The Bethe-Hessian down to the Percolation Threshold
The Bethe-Hessian is a symmetric matrix for which the negative spectrum has been observed to encode the informative structure of sparse stochastic block models. We prove that, in the stochastic block model where all vertices have expected degree $d>1$, the number of negative eigenvalues of the Bethe-Hessian is exactly the number predicted by the eigenvalues of the planted model lying outside the bulk spectrum. The condition $d>1$ is optimal, and matches a regime in which existing spectral approaches based on larger non-Hermitian matrices apply. Our result extends a theorem of Stephan and Zhu, who established the same conclusion under the assumption $d\geq 2$. Our improvement relies on two main ideas. First, we construct test vectors on the $2$-core, where degree fluctuations are substantially smaller, and then extend them to the entire graph while controlling the quadratic form. Second, we construct the test vectors using an isotropic basis of the underlying Markov random field, with coefficients adapted to each relevant planted eigenvalue. This allows us to control the fluctuations of the test vectors throughout the sparse regime.
2026-08-10
Random Width and Brightness: Polyhedral Density Theory, Reconstruction, and Gaussian Identifiability
Published • View Publication • BIB
Let U be uniformly distributed on the unit sphere. We develop a self-contained forward and inverse theory for the random width w_K(U) and brightness b_K(U) of three-dimensional convex bodies. For every full-dimensional polytope, a global spherical co-area formula expresses the width density as a finite sum of angular apertures determined by the normal fan of its difference body; in particular, the density is piecewise real analytic with a finite geometrically determined critical set. This theory yields exact densities for the width of the regular tetrahedron, resolving a question of Finch, and for the regular truncated octahedron, together with the tetrahedral brightness law and the equivalent rhombic-dodecahedral width law. On the inverse side, second- and third-order polarized cosine-transform moments reconstruct finite labelled direction systems whenever the observed triangles span the cycle space of the correlation graph; signed-graph switching describes the unavoidable ambiguity. In contrast, equal three-dimensional intrinsic volumes do not determine either the width law or the brightness law, even for centrally symmetric bodies. Removing the spatial rank constraint gives a dimension-free identifiability theorem for centered multivariate folded-normal vectors: pairwise absolute moments and an anchored family of triple absolute moments, comprising |m - 1|^2 labelled observations for a complete correlation graph, determine the correlation matrix up to diagonal sign conjugacy without fourth-order moments. A harmonic decomposition further identifies the degree-two variance contribution as a constant multiple of the squared Frobenius norm of the traceless part of the weighted frame operator and explains why this contribution vanishes under irreducible symmetry.
2026-08-04
Ranked spreadness and sample-based testing
In this note, we introduce the notion of ranked spreadness, a strengthening of the usual spread condition in which the elements of each member can be ordered so that their one-coordinate marginals decay geometrically with their rank. This additional structure removes the dependence on the maximum set size in random-containment estimates. We prove width-free hitting and weighted-concentration theorems for ranked-spread set systems, together with an elementary kernel-extraction theorem showing that ranked spreadness arises naturally in arbitrary distributions on small sets. Our main application is to the simulation of nonadaptive property testers by sample-based testers. If a one-sided tester has average query complexity $d$ and rejects every far input with probability at least $δ$, then, for every integer $c>d/δ$, it admits a one-sided sample-based simulation with expected sample complexity $O_{d,δ,|Σ|}\bigl(n^{1-1/c}\bigr)$. More generally, if positive inputs are rejected with probability at most $γ$ and far inputs with probability at least $δ>γ$, the same conclusion holds for every $c>d/(δ-γ)$. In particular, for constant-query nonadaptive testers we obtain an exponent $1-Θ(1/q)$, matching, up to the dependence on the rejection gap, the exponent conjectured by Fischer, Lachish, and Vasudev.
2026-07-21
Maximum Likelihood Estimation on the Grassmannian of Lines
We study the positive Grassmannian through the lens of algebraic statistics. A closed formula is presented for the maximum likelihood degree of the Grassmannian of lines. We conjecture that the probability simplex contains a unique local maximum, and we present computational evidence for this.
2026-07-10
A divisibility theorem for odd $J$-characteristics of two-level designs
We prove a divisibility theorem for the signed $J$-characteristics of two-level designs: if the number of factors $n$ is odd and every $J$-characteristic of a proper odd-cardinality subset of factors vanishes, then the top $J$-characteristic is divisible by $2^{n-1}$. As an arithmetic consequence, any two-level design whose $J$-characteristics vanish in orders one, two, three, five, and seven but which has a nonzero odd-order $J$-characteristic must have at least $256$ runs. This settles, uniformly in the number of factors, a conjecture of Eendebak, Schoen, Vazquez, and Goos (2023) on the nonexistence of certain strength-three even--odd designs with $56$ or $64$ runs. The divisibility bound is sharp at every odd order and is attained by the even-weight half-fraction.
2026-06-20
Recursive lower bounds for uniform set systems of bounded VC-dimension
For integers $n\ge d+1$, let $\mathsf{M}_d(n)$ denote the maximum size of a $(d+1)$-uniform family on an $n$-element ground set with VC-dimension at most $d$. For $n\ge2d+2$, the classical construction of Ahlswede and Khachatrian, later generalized by Mubayi and Zhao, gives \[ \mathsf{M}_d(n)\ge \binom{n-1}{d}+\binom{n-4}{d-2}. \] We introduce a two-cover lifting construction and prove the recursive lower bound \[ \mathsf{M}_d(n)\ge \binom{n-1}{d}+\binom{n-4}{d-2}+\mathsf{M}_{d-3}(n-5) \] for every $d\ge 3$ and $n\ge d+3$. Consequently, \[ \mathsf{M}_d(n)\ge \binom{n-1}{d}+\binom{n-4}{d-2}+\binom{n-6}{d-3}. \] Thus the Mubayi--Zhao conjecture on the exact value of $\mathsf{M}_d(n)$ for $n\ge2(d+2)$ is false for any $d\ge 3$. The proof is elementary and proceeds entirely through an explicit analysis of traces.
Tight $L_\infty$ Sample Complexity for Low-Degree and Sparse Boolean Polynomials
Motivated by the optimization of bounded binary black-box functions, we study the problem of learning polynomial surrogates over the Boolean hypercube. To ensure that optimizing the surrogate yields good solutions for the underlying objective, we require uniform $L_\infty$-error guarantees rather than the usual $L_2$-type guarantees. We characterize the minimax sample complexity of uniform estimation under subgaussian noise for two classes of bounded polynomials. First, for polynomials of degree at most $d$ on $n$ variables, the sample complexity scales as $n^{d+1}$. Second, for $s$-sparse Fourier-Walsh polynomials with $s \leq n$, it scales as $ns^2$. These rates differ structurally from the noiseless setting, where uniform exact recovery scales as $n^d$ and $ns$, respectively. Our lower bounds hold even for arbitrary adaptive learners, showing that the additional factors are intrinsic to the noisy cases. Standard Fourier-analysis tools for the $L_2$-norm do not naturally extend to the $L_\infty$-setting in a way that yields uniform guarantees. Our proofs overcome this difficulty by relying on suitably chosen auxiliary norms that serve as proxies for controlling the $L_\infty$-error. Together, our results provide a tight characterization of the sample complexity of learning optimization-safe polynomial surrogates.
2026-06-15
Euler Stratifications of Second Hypersimplices via Delta-matroids
We study Euler characteristics of scaled toric varieties arising from second hypersimplices. In algebraic statistics, these are closely connected to maximum likelihood (ML) degrees of toric models. We establish a correspondence between delta-matroids and the non-vanishing factors of the principal $A$-determinant, providing an explicit connection between delta-matroid theory and algebraic statistics. Using this framework, we show that a conjectured minimum ML degree is realizable by a suitable embedding of the variety. Furthermore, for second hypersimplices up to order six, we prove that this value is minimal among all embeddings, as conjectured by Clarke et al. (2024).
Sharp Low-Degree Thresholds for Planted-vs-Planted Testing
We establish the first sharp thresholds for low-degree polynomial tests in planted-vs-planted settings, where the goal is to determine with vanishing error which of two structured planted mechanisms generated the observed data. We prove matching low-degree upper and lower bounds for counting communities in the planted submatrix and planted dense subgraph models. The resulting testing threshold coincides, down to the sharp constant, with the known low-degree recovery threshold. In contrast, the task of weak testing, where the goal is to outperform random guessing, does not have a sharp threshold but rather a smooth transition, which we identify. To prove our results, we develop a framework for planted-vs-planted testing that builds on a latent-variable expansion originating in low-degree recovery and employs new methods to identify and prune non-signal contributions.
2026-06-01
Transitivity in Inhomogeneous Random Tournaments
Paired-comparison data are naturally represented by tournaments, where transitivity corresponds to the existence of a global ranking consistent with all pairwise outcomes. Accordingly, the classical Kendall-Smith coefficient of consistency measures deviations from transitivity in a tournament by counting the number of circular triads (directed $3$-cycles). In this paper, we characterize the fluctuations of the number of circular triads in inhomogeneous random tournaments and develop an inferential framework for the consistency coefficient. Specifically, we consider the $W$-random tournament model, where the comparison probabilities are determined by a tournamenton $W$, the analogue of a graphon in the tournament setting. We show that, for a $W$-random tournament on $n$ vertices, the number of circular triads exhibits three different fluctuation regimes, determined by suitable notions of regularity and uniformity of $W$. We further develop a novel tournamenton multiplier bootstrap that consistently approximates the limiting distribution of the circular-triad count in the relevant asymptotic regime. Combining this with procedures for testing regularity and uniformity, we design an algorithm for constructing confidence intervals for the consistency coefficient that is asymptotically valid for all tournamentons. We also obtain structural characterizations of tournamentons for which the limiting distribution of the number of circular triads exhibits specific degeneracies. These results can also be viewed through the lens of tournament quasirandomness and may be of independent interest.
2026-05-15
Bounds on the Number of Modes of a Gaussian Mixture Density
We derive explicit upper bounds for the number of nondegenerate critical points of a $k$-component Gaussian mixture density in $\mathbb{R}^d$, and the number of modes when the modal set is finite, together with lower bounds. By normalizing the critical-point equations by a reference component, for $k\ge2$ we get the direct Pfaffian bound \[ U_{\mathrm{het}}(d,k)=2^{\,d+\binom{k-1}{2}}\left(d+2\min(d,k-1)+1\right)^{k-1}. \] For the same parameter range, an exact elimination augmented by an algebraic reciprocal variable gives the alternative bound \[ U_{\mathrm{aug}}(d,k)= 2^{\binom{k-1}{2}}(d+1)\left((2k-1)d+2k-1\right)^{k-1}. \] Thus, for $k\ge2$, the best critical-point bound is their minimum. A Morse-theoretic argument improves the corresponding finite-mode upper bound to \[ \left\lfloor \frac{\min\{U_{\mathrm{het}}(d,k),U_{\mathrm{aug}}(d,k)\}+1}{2}\right\rfloor. \] In the homoscedastic case, for $k\ge2$, the direct bound improves to \[ U_{\mathrm{hom}}(d,k)=2^{\,d+\binom{k-1}{2}}\left(d+\min(d,k-1)+1\right)^{k-1}, \] an affine-rank reduction replaces $d$ by the affine rank of the component means, and an augmented homoscedastic reduction gives the dimension-free bound \[ U_{\mathrm{aug,hom}}(k)=2^{\binom{k-1}{2}+1}(2k)^{k-1}. \] On the lower-bound side, for $d,k\ge 2$ we obtain \[ L_{\mathrm{bin}}(d,k)=k+\max_{2\le r\le \min(d,k)}\binom{k}{r}, \] together with a padding-product family that in particular implies the linear lower bound $d+k-1$, and a seed-closure principle that packages product and padding constructions. We further give explicit bounds for the number of connected components of the critical set.
The stochastic block model has the overlap graph property for modularity
The overlap gap property (OGP) is a statement about the geometry of near-optimal solutions. Exhibiting OGP implies failure of a class of local algorithms; and has been observed to coincide with conjectured algorithmic limits in problems with statistical computational gap. We consider the Stochastic Block Model (SBM), where the graph has a planted partition with $k$ equal-size blocks which form the `communities', and where, for parameters $p>q$, vertices within the same community connect with probability $p$, while vertices in different communities connect with probability $q$, independently across pairs of vertices. Modularity--based clustering algorithms have become ubiquitous in applications. This article studies theoretical limits of local algorithms based on the modularity score on the SBM. We establish that modularity exhibits OGP on the SBM. This rules out a class of local algorithms based on modularity for recovery in the SBM, and shows slow mixing time for a related Markov Chain. Theoretically this is one of the few instances where OGP has been established for a `planted' model, as most such analyses to date consider the `null' model. As part of our analysis, we extend a result by Bickel and Chen 2009, who established that with high probability, the modularity optimal partition of SBM is $o(n)$ local moves away from the planted partition, where $n$ is the graph size. We show that, with high probability, any partition with modularity score sufficiently near the optimal value is close to the planted partition.
Thinned Quantile Shares are Universally Feasible
Quantile shares, introduced by Babichenko, Feldman, Holzman, and Narayan [STOC 2024], offer an ordinal, self-maximizing, and interpretable benchmark for fair division of indivisible goods, but their universal feasibility is known only conditional on the rainbow Erdős matching conjecture (EMC). Specifically, Babichenko et al. showed that assuming the rainbow EMC in the near-perfect matching regime, the $(1/2e)$-quantile share is universally feasible. In contrast, a simple argument shows that the $q$-quantile share can be infeasible for any $q > 1/e$. We introduce a one-parameter refinement of quantile shares, the $c$-thinned quantile share, obtained by thinning the inclusion probability in the random benchmark bundle by a factor of $c$ for a fixed constant $c\in(0,1]$. Our main result is that there exists a universal constant $c >0$ for which the $c$-thinned $e^{-c}$-quantile share is unconditionally universally feasible; this is best possible in the sense that for any $c \in (0,1]$, the $c$-thinned $q$-quantile share can be infeasible for any $q > e^{-c}$. Prior to this work, the only nontrivial share known to be universally feasible was Feige's residual maximin share. The thinning viewpoint also lets us remove the factor-two loss in the conditional result for the original quantile share: assuming the rainbow EMC, the $(1/e)$-quantile share is universally feasible.
2026-04-28
Exact Closed-Form Formulae for Linear and Circular Continuous Scan Statistics: $P_c(N - 1; N, w)$, $P_c(3; N, w)$, and $P(3; N, w)$
The continuous linear $P(k; N, w)$ and circular scan statistics $P_c(k; N, w)$ are fundamental tools in probability and spatial statistics, frequently used to detect clustering in uniform data. Let $X_1, X_2, \dots, X_N$ be independently and uniformly distributed random variables on a unit interval or unit ring. The exact distribution of these scan statistics relies on the minimum window width required to capture exactly $k$ points. Furthermore, the survival function $1 - P_c(k; N, w)$ directly corresponds to the geometric probability that if $N$ arcs of length $1 - w$ are uniformly and randomly placed on a unit circle, every point on the circle is covered at least $N + 1 - k$ times. Historically, evaluating the exact cumulative distribution functions, $P(k; N, w)$ and $P_c(k; N, w)$, relies heavily on complex recursive approximations. In this paper, we bypass these traditional recursive methods to derive direct, generalized closed-form expressions for some linear and circular continuous scan statistics. Specifically, we present the exact analytical solutions for $P_c(N - 1; N, w)$, $P_c(3; N, w)$, and $P(3; N, w)$ for arbitrary values of $N$ and window width $w$. These newly derived closed-form expressions not only provide exact baseline distributions for extreme spacings but also significantly simplify computational complexity compared to existing iterative approaches.