cs.IT ↗ arXiv
213 papers in this category
Elementary Symmetric Polynomial Inequalities for Centered Vectors and Matrices
We prove new inequalities for elementary symmetric polynomials (ESPs) for vectors that sum to zero, and for square matrices with zero row and column sums. We apply these results to obtain a unified upper bound on the mean-field approximation guarantee for permutation mixtures, as well as a sharp $χ^2$ version of the de Finetti theorem for finite sequences over a small alphabet. The main proof ideas were developed by the GPT-5.5 Pro model.
Cyclotomy, External Difference Families, and Algebraic Manipulation Detection Codes
External difference families (EDFs) are closely related to algebraic manipulation detection (AMD) codes, a cryptographic primitive that protects messages against additive tampering by an adversary who cannot observe the transmitted codeword. We first study a cyclotomic construction of external partial difference families over finite fields and derive a necessary and sufficient condition under which the resulting families are EDFs. For even block sizes up to $14$, we obtain explicit criteria in terms of quadratic and biquadratic residues, yielding new $R$-optimal weak AMD codes. We then construct bounded generalized strong external difference families using cyclotomic classes over finite fields, direct products of finite fields, and generalized cyclotomic classes over integer residue rings. These constructions give infinite families of systematic $G$-optimal strong AMD codes with flexible parameters.
On the Maximality of Additive Codes
An additive $(n,k,d)_{q^m/q}$-code is a $\mathrm{GF}(q)$-linear subspace of $\mathrm{GF}(q^m)^n$ of $\mathrm{GF}(q)$-dimension $km$ with minimum Hamming distance $d$. We first extend the Alderson--Bruen--Silverman (ABS) model of linear codes to the additive setting: a code of length $n$ with $q^{km}$ words over an alphabet of size $q^m$ admits an ABS model if and only if it is equivalent to a nondegenerate additive code. We then ask whether an additive code that admits an extension must admit an \emph{additive} extension. For linear codes ($m=1$) this is a theorem of Alderson and Gács. We characterize the additive codes admitting no additive extension as those whose associated projective system of flats is complete, and we prove that the answer to the question above is again affirmative for $(n,2,d)_{9/3}$-, $(n,2,d)_{4/2}$-, and $(n,3,d)_{4/2}$-codes. In contrast with the linear case, we show that the answer is negative in general. Scattered linear sets yield, for each square $q$, extendable additive $(n,2,d)_{q^2/q}$-codes admitting no additive extension. Further, a different method yields an extendable additive $(30,2,24)_{8/2}$-code with no additive extension. Consequently, for properly additive codes, completeness of the associated projective system does not imply maximality of the code. We conjecture that extendable $(n,2,d)_{p^2/p}$-codes, $p$ prime, always admit additive extensions.
Constrained Multi-Relational Graphons with Maximum Entropy
The principle of maximum entropy provides a fundamental framework for characterizing typical structures of large random networks subject to observable constraints. In their pioneering numerical experiments \cite{radin2014asymptotics}, Radin, Ren, and Sadun conjectured that entropy-maximizing graphons satisfying subgraph density constraints are stochastic block models a conjecture we term the RRS conjecture. While several special cases have been proven for single-relation graphs with specific constraint families, the general problem has remained open, particularly for multi-relational networks.
We resolve the RRS conjecture for constrained multi-relational graphons in the non-extremal regime, proving that entropy-maximizing solutions are step functions with finitely many blocks under the condition the subgraph density constraints are analytically independent and for almost all feasible combinations of sufficient statistics. Our proof employs a differential geometric technique to study solutions of constrained optimization problems in function space via functions with a finite parametrization (step functions). The two cornerstones of this work are: the generalization of subgraph density notion to $h$-subgraph density and the proof that manifolds that define the constrained region for the solutions maintain topological stability without developing new connected components under refinement. Together, these enable proving that no new global optima emerge in higher-dimensional spaces.
An Isodiametric Theorem and Lattice Diameter-Perfect Codes in $A_3$
The root lattice $A_n$, equipped with its graph distance (equivalently, one half of the ambient $\ell_1$ metric), is isometric to $\mathbb{Z}^n$ with the asymmetric Manhattan metric. We study two extremal problems in this space -- the isodiametric problem, i.e., determining the maximum anticode cardinality, and the (non)existence of linear diameter-perfect codes, i.e., lattice tilings by optimal anticodes -- and solve them in dimension $3$. We show that, for every integer $D\ge 0$, the largest cardinality of a diameter-$D$ subset of $A_3$ is $\binom{D+3}{3}+(D+1)\lfloor D^2/4\rfloor$, and this value is attained by the balanced difference of two discrete simplices. We then prove an integrality-refined simplex-packing obstruction: a sublattice of $\mathbb{Z}^n$ of asymmetric Manhattan distance greater than $D$ induces a lattice packing by $(D+1)Δ_n$ in $\mathbb{R}^n$. Combining this observation with the exact lattice-packing density of the tetrahedron yields a complete classification in dimension $3$: lattice diameter-perfect codes in $A_3$ exist precisely for $D=1$ and $D=2$. We also give the equivalent statement for perfect $B_h$ sets of cardinality four. Finally, we formulate a conjecture regarding optimal anticodes in arbitrary dimension, and restate it as an intersection problem for uniform multisets.
Improved lower bounds for the Shannon capacity of odd cycles
The Shannon capacity $Θ(G)$ of a graph $G$ quantifies the maximum rate at which information can be transmitted with zero error over a noisy channel. It is lower bounded by $α(G^d)^{1/d}$ for any $d$, where $α(G^d)$ is the independence number of the $d$-th strong power of $G$. We construct independent sets of size $134753$ in $C_7^{10}$, $21909$ in $C_{11}^{6}$, and $62530$ in $C_{13}^{6}$, improving the best known lower bounds for the Shannon capacity of these graphs to $Θ(C_7)\geq 134753^{1/10}>3.258020$, $Θ(C_{11})\geq 21909^{1/6}>5.289773$, and $Θ(C_{13})\geq 62530^{1/6}>6.300109$. We also improve the best known lower bounds on the independence numbers of several individual strong powers of odd cycles that do not improve the Shannon capacity lower bound. The constructions were discovered through iterative interactions with a Large Language Model (LLM), illustrating the potential of LLMs for finding explicit combinatorial constructions.
New lower bounds for binary constant-weight codes: $A(23,6,10)\geq 2979$ and $A(24,6,10)\geq 4214$
Let $A(n,d,w)$ denote the maximum size of a binary constant-weight code of length $n$, minimum distance $d$, and weight $w$. We construct explicit codes proving $A(23,6,10)\ge 2979$ and $A(24,6,10)\ge 4214$. These improve the best surviving explicit codes of sizes 2969 and 4174 and surpass the corresponding 1990 bounds 2970 and 4200 of Brouwer, Shearer, Sloane and Smith, whose code listings were lost. We also obtain $A(23,6,11)\ge 3539$ and $A(24,6,8)\ge 1855$. All four bounds are now listed in Brouwer's online table. The constructions use a coordinate decomposition in which one half is fixed to a known code and the complementary half is selected from its full cross-compatible pool using CHILS for maximum-weight independent set. For the 2969-word $A(23,6,10)$ incumbent, exact computations with two solver families prove insertion maximality and exclude every improving exchange deleting at most three codewords. We also analyze codes invariant under prime-order permutations: several cycle types are excluded exactly, the $5+1^{18}$ type has upper bound 499, and reproducible heuristic saturation evidence is reported for the remaining types, with $13+1^{10}$ left open. Code files, an independent validator, model descriptions, and computational logs are released.
Bounds and Limitations on Codes Achieving List Recovery Capacity
In coding theory, list recoverability is a fundamental concept which robustly captures how ``spread-out'' codewords are in a code. More formally, given a code $C \subseteq Σ^n$ and input lists $S_1, \dots, S_n \subseteq Σ$ of size at most $\ell$, list recoverability requires that there are at most $L$ codewords $c \in C$ such that $c_i \in S_i$ for at least $(1-ρ)n$ choices of $i \in [n]$. List recovery is an important question which has found applications in many areas, including complexity theory, property testing, compressed sensing, streaming algorithms, and cryptography.
As our first main result, we establish a tight ``generalized singleton bound''. Formally, we show that for constant $\ell, L,ρ$ and sufficiently large alphabets $Σ$, if we define $R^*=\frac{L+1-\ell}{L}-\frac{L+1}{L}ρ$, it is possible for a $(ρ,\ell,L)$ list-recoverable code to have rate $R^*-ε$ but impossible to have rate $R^*+ε$. One direction of our result already directly generalizes and improves a weaker impossibility result due to Goldberg, Shangguan, and Tamo.
For our second main result, we prove that there is a fundamental shortcoming in existing methods that aim to construct explicit, optimal list-recoverable codes. Indeed, recent work has constructed explicit codes achieving list-decoding capacity (along with other related properties) using a framework introduced in the work of Alon--Edmonds--Luby (AEL). We give a meta-analysis of such constructions by presenting an ``AEL framework'' which captures all such recent constructions in the literature. Within this framework, we show that no AEL-based code can break a recently-identified list-recovery barrier for additive and linear codes.
Cofilling Shattering: A Syndrome-Support Hierarchy for Check Erasures
Let $A:\mathbb{F}_2^n\to\mathbb{F}_2^m$ be a binary linear map with fixed coordinate bases, let $C_A=\ker A$, and let $λ_A(y)$ be the minimum Hamming weight of a preimage of the syndrome $y$. We define $\operatorname{Shat}_{q,s}(A)$ as the least common check support of a $q$-dimensional syndrome subspace whose every nonzero element has coset-leader weight at least $s$. It therefore distinguishes release of $q$ independent syndromes from release of a subspace with no easy linear combination. Deleting check coordinates $F$ releases $\ker A_{\bar{F}}/\ker A$, canonically isomorphic to $(\operatorname{im} A)[F]$.
Finiteness implies $R_q(C_A)\ge \mathsf{N}_2(q,s)$, where $\mathsf{N}_2(q,s)$ is the shortest length of a binary code of dimension $q$ and distance at least $s$; profile-Griesmer bounds independently control common check support. The hierarchy is coordinate-relabeling invariant but can change under a change of check basis. For the pair-repetition code $C_n=\{(x,x):x\in\mathbb{F}_2^n\}$, the standard realization $H_0=[I_n\ I_n]$ has $\operatorname{Shat}_{q,s}(H_0)=\mathsf{N}_2(q,s)$ whenever feasible. For every $q\ge 1$ and $s\ge 2$, with $n=\mathsf{N}_2(q,s)$, a row-equivalent realization of the same code has value $q$.
For a simplicial coboundary map $A=δ_k$, check erasure is top-face erasure and the released quotient is emergent cohomology. At $s=1$ the hierarchy reduces to generalized Hamming weights and is Tutte-determined; for $s\ge 2$, even identical labeled cut codes can have different values.
Revisiting the Stability of the Ingleton Inequality: A Tropicalization-Free Approach
The classical Ingleton inequality is known to hold for entropic points under specific exact conditional independence constraints. Recently, Matveev and Romashchenko (2026) investigated the stability of these implications, quantifying the extent to which the Ingleton inequality can be violated when a group of conditional mutual information terms is small but non-zero. While their proofs relied fundamentally on the complex framework of tropical probability spaces, we revisit these stability results using a completely tropicalization-free approach. By developing an alternative framework, we significantly streamline the underlying concepts and proofs, derive explicit error terms, and improve some estimates. Furthermore, we resolve an open problem posed in prior work by exhibiting a new infinite family of entropy inequalities that establishes the stability of the sum of two Ingleton expressions.
Decoding Desarguesian spread codes beyond half minimum distance
Spread codes are a well-known family of constant-dimension subspace-metric codes. For constant dimension $k$ and ambient space dimension $n$ being a multiple of $k$, these codes have minimum distance $2k$ and a rich geometric structure. In this paper, we study the decoding capabilities of the Nearest Neighbor Decoder for Desarguesian spread codes, establishing that unique decoding is still achievable beyond half the minimum distance. Motivated by this, we develop a new decoding algorithm to uniquely decode Desarguesian spread codes in the presence of both insertions and deletions, which increase and decrease, respectively, the dimension of the transmitted codeword. Even when the sum of the dimensions of insertions and deletions exceeds half the minimum distance, provided that deletions are of dimension at most $k-2$, the algorithm succeeds with a small decoding failure. We also propose two refinements to this algorithm that, empirically, can handle nearly as many insertions as the Nearest Neighbor Decoder.
Independent Sets in Multiset Profile Graphs via Weighted Local Covers
Let $G_q(d)$ be the unit-transfer graph on the nonnegative integer vectors whose $q$ coordinates sum to $d$, equivalently on the multiplicity profiles of size-$d$ multisets over $q$ symbols. The prime-checksum conjecture predicts that, for prime $q$ and all sufficiently large $d$, a largest independent set is a fiber of the natural cyclic checksum. We introduce weighted local covers of $G_q(d)$ by translated induced subgraphs. For fixed $q$, capped anchor profiles reduce the covering conditions for infinitely many degrees to a finite rational linear system.
This method gives new proofs of the known cases $q=3$ and $q=4$ and determines $α(G_q(d))$ exactly for $q=5$ and $q=7$ in every degree, thereby proving the next two odd-prime cases of the conjecture. In the complementary regime where $d$ is fixed and $q$ grows, a partition-orbit reduction solves degree five for $q\ge7$, gives exact power-of-two families in degrees six, eight, and ten, and yields an asymptotically sharp upper bound through three terms for every fixed $d\ge7$. All computer-assisted assertions reduce to finite rational or integer systems and are supported by independently checkable certificates.
Construction of Generalized Weighing-Hadamard Matrices over Finite Fields
The existence, several properties, and constructions of Generalized Weighing-Hadamard (GWH) matrices over finite fields are addressed in this work. We study the subset of invertible GWH matrices and show that it forms a group under matrix multiplication. Besides that, we introduce a strong notion of equivalence between such matrices, defined via orthogonal transformations, and further prove that the corresponding quotient group by the subgroup of orthogonal matrices is abelian. Finally, we discuss some applications of these matrices in coding theory
Minimum distance and decoding of Coxeter codes
A binary Coxeter code associated with a finite Coxeter system $(W,S)$ is an ${\mathbb F}_2$-linear span of indicators of standard cosets of a fixed rank. Coxeter codes, introduced in a recent paper by N. Coble and A. Barg, are a generalization of Reed--Muller codes which arise when $W={\mathbb Z}_2^m$ is the Coxeter group of type $mA_1$. In that paper, the authors proposed a conjectural value for the minimum distance of a general Coxeter code. This conjecture is proved in the present work. As a consequence, we obtain a Coxeter-theoretic generalization of Reed's majority-logic decoding algorithm for Reed--Muller codes.
On the Etzion-Silberstein conjecture for block Ferrers diagrams
Ferrers diagram rank-metric codes are rank-metric codes with prescribed support, and their dimension is bounded from above by the Etzion--Silberstein bound. In this paper, we study this problem for block Ferrers diagrams, namely Ferrers diagrams whose dots are grouped into square blocks of a fixed size. Motivated by the diagonal construction for MDS-constructible Ferrers diagrams, we introduce the notion of MSRD-constructibility, where MDS codes on diagonals are replaced by maximum sum-rank distance (MSRD) codes on block diagonals. We show that MSRD-constructible pairs yield optimal Ferrers diagram rank-metric codes over sufficiently large finite fields. We then relate MSRD-constructibility of a block Ferrers diagram to MDS-constructibility of its contraction, proving an equivalence when the distance is compatible with the block size and giving lifting criteria in the general case. As a consequence, we obtain MSRD-constructibility for strictly block-monotone and initially block-convex diagrams. Finally, we prove a reduction to block triangular diagrams and use it to obtain new arbitrary-field cases of the Etzion--Silberstein conjecture for MSRD-constructible block Ferrers diagrams.
Combinatorial constructions of Schubert subspace codes
We study Schubert subspace codes, which are constant-dimension subspace codes with prescribed intersection conditions with a fixed subspace. Our goal is to construct codes of maximum possible size in the extremal distance cases where a natural counting upper bound applies. We give two families of constructions. The first one uses a direct-sum decomposition of the ambient space, together with partial spreads and colorings of powers of $q$-Johnson graphs. For this construction, we also prove necessary conditions, which show how chromatic and clique obstructions arise. The second family is obtained by field reduction from evasive and scattered subspaces over extension fields. This gives codes whose size can be computed exactly in the scattered case and recovers the only previously known construction as a special case.
Unique Insertion Error Patterns in Levenshtein's Reconstruction Problem
Levenshtein's sequence reconstruction model plays an essential role in information retrieval of advanced memory systems, such as the DNA-based storage systems. In the model, a word $\mathbf{x}\in\mathbb{Z}_q^n$ is transmitted through $N$ noisy channels, and the goal is to recover it. Errors occurring in the channels usually involve substitutions, insertions and deletions. Our focus is on insertions. One of the main questions in this context is determining the minimum number of channels $N$ required to recover the transmitted word $\mathbf{x}$. The original formulation of the reconstruction problem requires that all the output words from the channels are distinct. However, different insertion errors may lead to the same output words. In this paper, we investigate two reconstruction models where the channels are allowed to produce identical output words even though different insertion errors occur in the channels. These two models, called \textit{the multiset model} and \textit{non-multiset model}, generalize the Levenshtein's model. We denote the minimum number of channels required to \textit{unambiguously} recover the transmitted word $\mathbf{x}\in\mathbb{Z}_q^n$ by $N_q^m(n,t)+1$ in the multiset model and $N_q^{nm}(n,t)+1$ in the non-multiset model, where $t$ is the exact number of insertions occurring in a channel. We determine $N_q^m(n,1)$ and $N_q^{nm}(n,1)$ for all $n$ and $q$, and show the somewhat surprising fact that $N_q^m(n,1)=N_q^{nm}(n,1)$. We also provide a full characterization of the words attaining this value and give a general lower bound on $N_q^m(n,t)$ for $t\ge1$ and a recursive upper bound. For $t=1$, we construct codes $C'\subseteq\mathbb{Z}_q^{n+2}$ from codes $C\subseteq\mathbb{Z}_q^n$ such that the number of channels required to determine the transmitted word $\mathbf{x}\in C'$ is small. This construction is shown to be optimal for certain parameters.
Minimum distances of LDPC codes in 5G standard
We propose several approaches for bounding the minim\-um distances of the family of quasi-cyclic LDPC codes in the 5G NR standard. In particular, we show that the high-rate [9984, 8448] and the low-rate [25344, 8448] BG1 5G LDPC codes have minimum distances in the ranges {8..14} and {22..57}, respectively. Also we propose a new early termination approach based on circulant modular reduction, which significantly lowers syndrome calculation complexity for the LDPC decoder.
Function-Counting Theory for Low-Dimensional Data Structures
The success of deep learning models in classification and regression is widely attributed to the low-dimensional structure that real-world data tend to exhibit, despite their high-dimensional representation. This work attempts to provide a mathematical framework for binary classification on low-dimensional data, building on Cover's (1965) function-counting theory. With our framework, we aim to address the question of how the low-dimensional structure of the data affects the classification capabilities of learning models. Cover's theory relies on a general position assumption that blinds it to the underlying data structure. We refine this assumption to account for the low-dimensionality of the data and derive dichotomy counts that reflect the data structure. We further extend Cover's separation capacity and problem of generalization to the low-dimensional setting, enabling the impact of the underlying data structure on both to be analyzed.
Guesswork Under Linear Constraints: Exact Exponent for Coset Decoding
We establish the exact exponential growth rate of the $ρ$-th moment of the constrained guesswork $G_{\mathrm{coset}}$ -- the rank of the true noise vector within its syndrome coset of a random binary linear code under i.i.d.\ Bernoulli$(p)$ noise: \( \lim_{n\to\infty} \frac{1}{n}\log_2\Eb\!\left[G_{\mathrm{coset}}^ρ\right] = ρ\,h_{\frac{1}{1+ρ}}(p)\;+\;ρ(R-1), \, ρ>0, \) where $h_α(p)$ is the binary Rényi entropy and $R=k/n$ is the code rate. The exponent shifts down by exactly $ρ(1-R)$ relative to the unconstrained Arıkan--Merhav exponent, with each of the $n(1-R)$ parity checks contributing equally. Finite-length simulations confirm convergence from below. We further establish: (i)~a transfer theorem expressing the partition-function exponent in terms of an arbitrary weight-enumerator growth rate $g(δ)$; (ii)~the exact exponent for $L_n$-list (``$k$-th'') constrained guesswork; and (iii)~a sharp second-order refinement of order $ρ\log_2 n$. Beyond the binary i.i.d.\ setting, we prove a universality theorem: for any code ensemble $\mathcal{E}$ whose weight enumerator concentrates at rate $g_{\mathcal{E}}(δ)$, the guesswork exponent equals $(1+ρ)ψ_{1/(1+ρ)}(g_{\mathcal{E}})-ρ\,ψ_1(g_{\mathcal{E}})$, where $ψ_α(g)=\sup_δ[g(δ)+α\ell(δ)]$. As concrete applications, we instantiate this theorem for the $q$-ary extension, $Λ_q(ρ)=ρ\,h^{(q)}_{1/(1+ρ)}(P)+ρ(R-1)\log_2 q$, and for Gallager's regular LDPC ensemble, obtaining a closed-form guesswork exponent via an exact finite-length identity for the ensemble-average weight enumerator.