cs.IT ↗ arXiv
213 papers in this category
Multi-View Block Distance Distributions and Linear Programming Bounds for Locally Recoverable Codes with Availability
We develop a multi-view linear-programming framework for locally recoverable codes with arbitrary fixed availability $a$. For any retained order $1 \le s \le a$, the selected helper sets, the recovered coordinate, and their complement form an $(s+2)$-part partition. Recording the Hamming distance on all blocks preserves both compatibility among the selected repair alternatives and their coupling with the remaining coordinates. The resulting joint distribution satisfies centered counting identities, product-Krawtchouk positivity, and collision inequalities from the local-distance condition; for fixed $s$, these constraints give a polynomial-size relaxation for arbitrary, possibly nonlinear, codes over any finite field. Retaining one view recovers the three-block model of our companion paper. We develop the first genuinely multi-view case, $s=2$, in detail and specialize the exact computations to availability $a=2$. The resulting four-block LP projects to both the one-view three-block model and a globally conditioned two-view relaxation, making explicit the information lost by each coarsening. Exact rational primal-dual certificates together with checked constructions prove $M_{\max}(2,8,4,2,2,2)=8$, $M_{\max}(3,7,3,2,2,2)=27$, and $M_{\max}(4,7,3,2,2,2)=64$; in all three cases the four-block bound is strictly stronger than both coarsenings.
The half-rate linear programming bound for binary codes is $\frac12-\frac1π$
In their work on sphere packing and the modular bootstrap, Afkhami-Jeddi, Cohn, Hartman, de Laat, and Tajdini conjectured the exact high-dimensional exponent of the Cohn-Elkies sphere-packing linear program. OpenAI's Chapter 1 subsequently proved their conjecture by establishing that both Fourier sign-uncertainty radii are $(1/π+o(1))\sqrt d$. We prove the binary coding analogue: \[
R_D\!\left(\frac12-\frac1π\right)=\frac12. \] We also formulate the two Krawtchouk sign-uncertainty problems and determine both of their asymptotics. If $A^{\mathrm K}_{\pm}(n)$ denotes the smallest radius $r$ for which a nonzero Krawtchouk $(\pm1)$-eigenfunction $f$ exists with $f(0) = 0$ and $f(x) \ge 0$ for all $|x| \ge r$, then \[
\frac{A^{\mathrm K}_{\pm}(n)}n\longrightarrow
\frac12-\frac1π. \] The lower bound proves a mass-concentration principle similar to OpenAI's Chapter 1 for Hamming space.
The upper bound, on the other hand, follows the approach of the spherical-code construction in OpenAI's Chapter 2. Gay, Jeronimo, and Liu formulated a hierarchy for binary codes analogous to the spherical-code construction and improved the best known binary coding rate bounds by evaluating the first level of the corresponding hierarchy. We prove this hierarchy bounds the Delsarte program and give a construction at arbitrarily deep levels of the hierarchy, attaining the upper bound in the limit. The construction uses an $N$-qubit generalization of the pure-state channel of Alrabiah and Guruswami.
Binary Multiple-Node-Erasure-Correcting Codes over Complete Graphs: Constructions, q-Ary Metric Balls, and Duality
We study linear codes whose coordinates are the ordinary edges and self-loops of complete undirected graphs; a node erasure removes all coordinates incident with a failed vertex. The construction results are binary. For triple-node erasures, we extend the published cyclic construction by allowing a suitable cyclic check slope to depend on the prime graph length. An explicit determinant test proves that one of three fixed slope choices works at infinitely many prime lengths, unconditionally, and gives redundancy $3n-2$, one bit above the graph Singleton bound. We also give Singleton-optimal triple-node codes at $n=6,8,10,12$, together with a general ordinary-edge framework that isolates the remaining loop-completion problem. When $2$ is primitive modulo an odd prime $n$, a binary multi-slope construction corrects every $ρ$-node erasure for $2\leqρ<n$, with redundancy $ρn-(ρ-1)$ in the range $2\leqρ\leq(n+1)/2$. Returning to arbitrary prime powers, we derive exact generating transforms and inclusion--exclusion formulas for node-metric ball volumes, fixed-radius asymptotics, and packing, existence, and covering bounds. Finally, for the complementary clique-erasure metric, we obtain an exact weight enumerator and a Singleton-optimal node--clique duality.
Approximate Uniformity in Finite Convolution Models
The problem of the uniform law being a sum of two independent distributions has been well studied. Here, we study the approximation of the uniform law to a sum of two distributions with fixed support, under the following discrepancies: the Manhattan distance, the Euclidean distance, the forward Kullback--Leibler divergence and the Wasserstein-one distance based on the line metric. The problem of membership of the uniform law in this model has been well studied. Explicit results are obtained in two models, one where one of the distributions is a Bernoulli and the other when both distributions have the same support, including one conjecture. Finally, we show an application of our reconstruction method to a recent conjecture about the coding capacity of an additive noise channel.
A weak Hellinger inequality for noisy Boolean channels
A weak form of the Hellinger conjecture of Anantharam, Bogdanov, Chakrabarti, Jayram, and Nair for the binary symmetric channel is proved: dictator functions maximize Hellinger $Φ$-entropy among all Boolean functions of the input and all one-bit statistics of the output of a noisy channel. The technical heart of the matter is an explicit inequality in three real parameters, which is proved using explicit polynomial approximations and computer-assisted positivity checks. The results are also formally verified in Lean 4.
Improved upper bound on the number of distinct k-decks for any k and alphabet size by counting the independent parameters
Data stored in synthetic DNA is retrieved by shotgun sequencing, which returns short subsequences rather than the stored word itself. A natural abstraction of this readout is the $k$-deck of a word: the vector recording how often each word of length $k$ occurs as a subsequence. Two stored words are distinguishable from their readouts exactly when their $k$-decks differ, so the number $D_{q,k}(n)$ of distinct $k$-decks of words of length $n$ over an alphabet of size $q$ measures what a length-$k$ readout retains.
We analyse the degrees of freedom remaining in a $k$-deck once all shorter decks are fixed. Within each class of words having prescribed letter multiplicities, the length-$k$ entries are confined to an affine subspace whose dimension is exactly the number of Lyndon words with the same multiplicities, which we give in closed form as a Möbius sum. Writing $L_q(j)$ for the number of Lyndon words of length $j$ over an alphabet of size $q$, we deduce the improved upper bound \[
D_{q,k}(n)=O\!\left(n^{E_q(k)}\right),\qquad E_q(k)=\sum_{j=1}^{k}j\,L_q(j)-1 . \] In the case of a binary alphabet this bound satisfies $D_{2,k}(n)=O\!\left(n^{4\cdot 2^{k-1}}\right)$.
We then prove matching lower bounds in the first two nontrivial cases: $D_{q,2}(n)=Θ\!\left(n^{q^2-1}\right)$ for every alphabet size $q$, and $D_{2,3}(n)=Θ(n^{9})$ for the binary alphabet. The latter confirms, for $q=2$ and $k=3$, our conjecture that the upper bound has the correct degree for every $q$ and $k$.
On different notions related to APN mappings
An APN mapping $F:\mathbb{F}_{2^n}\to \mathbb{F}_{2^n}$ is a polynomial characterized by the non-vanishing property on 2-flats. In this work, we analyze notions that are closely related to this property. To understand which $k$-flats of $\mathbb{F}_{2^n}$ remain flats under $F$, we study the $k$-breaking. The function $x^{-1}$ has been studied in the past in this context---we extend this study to general mappings and characterize the 2-breaking of APN functions. Recently, two generalizations of the APN property have been introduced: $k$-strongly non-normality and $k$-th-order sum-freedom. Sum-freedom generalizes the non-vanishing property of APN functions to higher dimensional flats. We provide in-depth observations of the relations between the breaking property, strongly non-normality and sum-freedom. We introduce a fourth concept called $k$-strongly breaking, which implies the breaking property. We derive several structural results for both notions and give a characterization of a subclass of APN functions in terms of the 2-strongly breaking property. We propose a different perspective of the non-vanishing property via a natural character transformation, which is closely related to the sum-of-square indicator of the components of $F$. We derive a precise value for the total sum of the sum-of-square indicators of $F$. With this approach, we provide a simple answer to Open Problem 4 in IEEE Trans. Inf. Theory 52(9): 4160-4170, 2006. Moreover, it allows us to explore balancedness properties of polynomials, one of which characterizes component-wise APNness, for odd $n$, and provides a natural extension to any dimension. We show that Dillon's APN permutation and the Gold functions satisfy a related property, termed $k$-balanced, which is presented under our framework.
Discrepancy for Random Linear Codes
We show that random linear codes (RLCs) possess nearly optimal discrepancy-type properties in a broad range of settings. Our main results are two general discrepancy theorems: one controls all translates of a fixed test, and the other controls large families of Fourier-pseudorandom tests. Two motivating examples follow:
First, RLCs behave essentially like unstructured random codes for list-decoding from errors above capacity. More precisely, an RLC $C\subseteq \mathbb{F}_q^n$ of rate $1 - \frac{1}{n}\log_q|B_ρ| + \varepsilon$, where $|B_ρ|$ is the volume of a radius-$ρ$ Hamming ball in $\mathbb{F}_q^n$, satisfies $|C \cap B| = (1\pm o(1)) \frac{|C|\cdot |B|}{q^n}$ simultaneously for all radius-$ρ$ Hamming balls $B$ with high probability. This vastly generalizes the previously best known fact that RLCs of this rate have covering radius at most $ρn$ with high probability (Blinovsky, 1987).
Second, over prime fields, RLCs behave essentially like unstructured random codes for zero-error list-recovery, and list-recovery from erasures, above capacity. More precisely, for a prime $q>2$ and input list size $2\leq \ell\leq q-1$, an RLC $C\subseteq \mathbb{F}_q^n$ of rate $1-\log_q \ell+\varepsilon$ will satisfy $|C \cap S| = (1\pm o(1)) \frac{|C|\cdot \ell^n}{q^n}$ simultaneously for all combinatorial rectangles $S=S_1\times S_2\times\cdots\times S_n$, where $|S_i|=\ell$ for all $i$, with high probability. An analogous result also holds when we can bound $|S_i|$ only for some of the $i$'s.
We use this to show the abundance of locally leakage-resilient $n$-party linear ramp secret sharing schemes with any linear reconstruction threshold and sublinear threshold gap $O(n/\log n)$ over fields of polynomial size $q=Θ(n^γ)$ for a constant $γ\in(0,1/5)$. Prior work was stuck at reconstruction thresholds above $n/2$ for both threshold and ramp schemes.
LDPC Fractus Codes: Sparse Codes with Recursive Structure
We introduce a new family of recursively constructed sparse matrices, termed Fractus matrices, and investigate their use in constructing low-density parity-check (LDPC) codes. Generated through a self-similar recursive process, these matrices yield regular sparse parity-check matrices while preserving key structural properties across successive iterations. This recursive structure enables an efficient encoding algorithm with computational complexity that is nearly linear in the block length. Decoding is performed using standard iterative message-passing algorithms, thereby retaining the low-complexity decoding characteristic of LDPC codes. The proposed construction produces Tanner graphs with girth six and guarantees a minimum Hamming distance of at least $\ell+1$. We establish several algebraic properties of Fractus matrices, including sparsity, regularity, recursive decomposition, and symmetry under the flip-transpose operation. In addition, we show that the family of Fractus matrices admits a natural lattice structure and that the associated LDPC codes inherit corresponding lattice-theoretic properties. These results establish a connection between order theory and coding theory. Overall, the proposed framework integrates recursive matrix constructions, efficient encoding, graph-theoretic analysis, and lattice theory into a unified algebraic approach to the design and analysis of scalable LDPC codes.
A Hyperbolic Bound for File Retrieval in DNA-Based Data Storage
In DNA-based storage systems, data are retrieved by sequencing molecules sampled from a DNA pool. We study how the way coding redundancy is shared between files affects their expected retrieval times. Our focus lies on the case of two files that are encoded by a systematic linear code over an arbitrary finite field. We consider the conjecture that the sum obtained by dividing each file dimension by its expected retrieval time is at most one whenever the dimension of at least one file is more than one. For this, we develop a geometric view of the retrieval process. As molecules are sampled, we follow the growing span of the corresponding columns of the generator matrix and track how much of this span comes from each file. Among the samples that enlarge the overall span, this lets us compare those that make progress toward recovering both files with those that make progress toward neither. We call a column mixed if the corresponding encoded symbol combines information from both files. In this work, we sharpen a projection bound and use it to control the effect of mixed columns. We prove the conjecture whenever the total information dimension is at least twice the number of mixed columns plus two. For equal-sized files, this allows up to one fewer mixed column than the dimension of either file and extends the previous result for codes with no mixed columns.
Improved bounds for constant-power and low-power error-correcting cooling codes
Low-power error-correcting cooling (LPECC) codes and constant-power error-correcting cooling (CPECC) codes provide error correction while controlling power consumption and thermal effects in on-chip buses. In this paper, we study binary CPECC and LPECC codes with \(e=w-3\). For CPECC codes, we extend the applicability of the upper bound previously obtained by Zhao and Zhang from the quadratic-order condition \(w\ge 2t(t+1)+2\) to \(w\ge w_0(t)\), where \(w_0(t)\sim \sqrt{2}\,t^{3/2}\). Using Steiner systems, we show that the CPECC bound is attainable and asymptotically tight for fixed \(t,w\). For LPECC codes, we establish the new upper bound \(\left\lfloor\frac{\binom{n+2}{3}}{\binom{w+t}{3}}\right\rfloor\) for \(w\ge μ(t)\), where \(μ(t)\sim \sqrt{2}\,t^{3/2}\). This bound is strictly smaller than the previous bound of Zhao and Zhang whenever both apply, and is asymptotically tight for fixed \(t,w\) in the stated range.
A proof of the generalized packing-covering conjecture
The generalized packing--covering conjecture of Elimelech, Firer and Schwartz asserts that, for every linear code $\mathcal{C}$ and every admissible order $t$, the $t$-th generalized Hamming weight $d_t(\mathcal{C})$ and the $t$-th generalized covering radius $R_t(\mathcal{C})$ satisfy $d_t(\mathcal{C})\le 2R_t(\mathcal{C})+2$. We give a computer-assisted proof of the conjecture for every linear code over every finite field and every admissible order. Combining a parity-check reformulation of the conjecture, bounds on the length of putative counterexamples, and successive puncturing arguments, we settle all orders $t\ge 32$ and reduce the remaining orders to finitely many parameter tuples, which we exclude by an exact computer verification.
A proof of the Braun-Etzion-Vardy bound for binary subspace codes
The notion of a linear subspace code in a projective space was introduced by Braun, Etzion and Vardy (2013), who conjectured that a linear subspace code in the projective space has at most 2^n codewords. We resolve this conjecture. A novel character-theoretic argument is used to prove the conjecture and the proof is self-contained.
A Continuous Projection Converse and a Weighted Construction for Uniquely Decodable Code Pairs
We study the maximum sum rate of uniquely decodable code pairs. A weighted refinement of a complement-gluing construction yields an explicit code of length 672 and sum rate exceeding $1.318639029203$. An iterated projection argument, combined with a conditional entropy estimate and a justified continuous limit, gives an upper bound of $1.480063425539$.
Sunflowers of Reed--Solomon Codes
We introduce and study Reed--Solomon sunflowers, namely families of Reed--Solomon codes whose pairwise intersections are all equal to the same fixed subspace. This notion lies at the intersection of extremal subspace combinatorics and coding theory: it can be viewed as a structured version of the sunflower problem in the Grassmannian, and it naturally produces constant-dimension subspace codes with prescribed minimum distance. We focus mainly on the case in which the center is the one-dimensional space generated by the all-one vector. We give an algebraic criterion, expressed in terms of generalized $V$-matrices, ensuring that a family of Reed--Solomon codes forms such a sunflower. We then study the size of these families through counting and constructions. In dimension two, we show that all distinct Reed--Solomon codes form a sunflower and determine its size by counting Reed--Solomon codes up to affine equivalence of their evaluation vectors. For fixed dimension $k\geq3$ and length $\ell\geq2k-1$, we give an explicit recursive construction with $Ω_{k,\ell}(q^{\lfloor\ell/(2k-1)\rfloor})$ petals and a greedy existence argument with $Ω_{k,\ell}(q^{\ell-2k+2})$ petals as $q\to\infty$. We also apply the greedy argument to obtain families of $[\ell,k]_q$ MDS codes of size $Ω_{k,\ell}(q^{2(\ell-2k+2)})$, whose pairwise intersections have dimension at most one but need not be equal.
Optimal and minimal $p$-ary linear codes from generalized order ideals of hierarchical posets
Hyun, Kim, Wu and Yue constructed optimal and minimal binary linear codes from order ideals of hierarchical posets with two levels. Two different generalizations of the underlying antichain (simplicial complex) setting to odd characteristic are known: down-sets of $\mathbb{F}_p^n$ under the componentwise order, and support-closed subsets of $\mathbb{F}_q^m$. No generalization of the poset setting itself has appeared. We introduce generalized order ideals of a poset of order $p-1$, obtained by attaching multiplicities in ${0,\dots,p-1}$ to the elements of a poset, and study the two natural notions of order ideal that arise for hierarchical posets with two levels. Whenever the ideal meets the upper level, the resulting defining sets are neither down-sets nor support-closed. We determine the weight distributions of the associated complement codes, exhibit a family of Griesmer codes in which the upper element carries an arbitrary multiplicity, and, via the characteristic function of a generalized order ideal, obtain an infinite family of minimal $p$-ary codes of length $p^n-1$ and dimension $n+1$ violating the Ashikhmin--Barg condition.
Quantitative tiling stability from quadratic discrepancy in Hamming spaces
Quadratic ball discrepancy measures how uniformly a code meets Hamming balls over all centers and radii. Perfect codes minimize this quantity among codes of the same cardinality. We prove a quantitative refinement: at the parameters of every nontrivial perfect code, excess discrepancy is at least a positive multiple of the discrepancy at the correction radius. The latter equals the squared deviation of the covering multiplicity from one, divided by the square of the code cardinality. The comparison applies to all codes of that cardinality. For one-error parameters, the coefficient is at least one, which is sharp uniformly over the parameters; explicit coefficients depending on the parameters can be much larger. For two-error parameters of length at least five, we obtain a positive coefficient against an algebraic benchmark under the sphere-packing and Lloyd integrality conditions, even when existence of a perfect code is unknown. The multiplicity defect also measures the failure of uniform ball noise to smooth a code distribution. Consequently, the same discrepancy excess gives explicit bounds on holes, overlaps, and divergence from the uniform distribution.
Perfect codes as exact minimizers of quadratic discrepancy in q-ary Hamming spaces
Total quadratic ball discrepancy measures the deviation of codeword counts from uniformity over all Hamming balls. We prove that, whenever a Hamming space and cardinality admit a perfect code, the discrepancy minimizers among all subsets of that cardinality are precisely the perfect codes. This extends Barg's binary minimizing result and characterizes equality over arbitrary finite alphabets. At the necessary parameter sets for correcting one or two errors, we give explicit lower bounds whose attainment is equivalent to perfect tiling, even when existence is unresolved. For alphabets of size at least four, the numerical values follow from known universal energy bounds; we establish their discrepancy normalization and the attainment criterion. The ternary discrepancy potential falls outside the completely monotonic regime of those bounds, and separate exact certificates settle the ternary Hamming and Golay cases.
The maximal hard-core model on the triangular lattice
The well-known hard-core model on the triangular lattice $\mathbb A$ is an interaction model defined on the independent sets of the lattice, parameterized by an activity parameter $λ> 0$. In this work, we consider an extension of this model to maximal independent sets, which we call the maximal hard-core model. We show that in the high-activity regime ($λ> 10^6$), the model admits at least two Gibbs measures on maximal independent sets in $\mathbb A$. In the low-activity regime ($λ$ close to $0$), we apply Pirogov-Sinai theory to characterize extremal periodic Gibbs measures of the model. Finally, we derive bounds on the capacity of a recoverable system on the lattice associated with the maximal hard-core model.
The maximal hard-core model as a recoverable system: Gibbs measures and phase coexistence
Published in Journal of Statistical Physics, 2026
• Search Publication
Recoverable systems provide coarse models of data storage on the two-dimensional square lattice, where each site reconstructs its value from neighboring sites according to a specified local rule. To study the typical behavior of recoverable patterns, this work introduces an interaction potential on the local recovery regions of the lattice, which defines a corresponding interaction model. We establish uniqueness of the Gibbs measure at high temperature and derive bounds on the entropy in the zero- and low-temperature regimes.
For the recovery rule under consideration, exactly recoverable configurations coincide with maximal independent sets of the grid. Relying on methods developed for the standard hard-core model, we show phase coexistence at high activity in the maximal case. Unlike the standard hard-core model, however, the maximal version admits nontrivial ground states even at low activity, and we manage to classify them explicitly. We further verify the Peierls condition for the associated contour model. Combined with Pirogov-Sinai theory, this shows that each ground state gives rise to an extremal Gibbs measure, proving phase coexistence at low activity.