formalized
129 papers tagged with this keyword
Composition Direction of Seymour's Theorem for Regular Matroids -- Formally Verified
Seymour's decomposition theorem is a hallmark result in matroid theory presenting a structural characterization of the class of regular matroids. Formalization of matroid theory faces many challenges, most importantly that only a limited number of notions and results have been implemented so far. In this work, we formalize the proof of the forward (composition) direction of Seymour's theorem for regular matroids. To this end, we develop a library in Lean 4 that implements definitions and results about totally unimodular matrices, vector matroids, their standard representations, regular matroids, and 1-, 2-, and 3-sums of matrices and binary matroids given by their standard representations. Using this framework, we formally state Seymour's decomposition theorem and implement a formally verified proof of the composition direction in the setting where the matroids have finite rank and may have infinite ground sets.
All rectangles exhibit canonical Ramsey property
In a seminal work, Cheng and Xu proved that for any positive integer \(r\), there exists an integer \(n_0\), independent of \(r\), such that every \(r\)-coloring of the \(n\)-dimensional Euclidean space \(\mathbb{E}^n\) with \(n \ge n_0\) contains either a monochromatic or a rainbow congruent copy of a square. This phenomenon of dimension-independence was later formalized as the canonical Ramsey property by Geheér, Sagdeev, and Tóth, who extended the result to all hypercubes, and to rectangles whose side lengths \(a\), \(b\) satisfy \((\frac{a}{b})^2\) is rational. They further posed the natural problem of whether every rectangle admits the canonical Ramsey property, regardless of the aspect ratio.
In this paper, we show that all rectangles exhibit the canonical Ramsey property, thereby completely resolving this open problem of Geheér, Sagdeev, and Tóth. Our proof introduces a new structural reduction that identifies product configurations with bounded color complexity, enabling the application of simplex Ramsey theorems and product Ramsey amplification to control arbitrary aspect ratios.
Towards a format for describing networks / 1. Networks and knowledge graphs
Published in Proceedings of Information Society 2025, SiKDD, 6-10 October 2025, Ljubljana, Slovenia, p. 102-105
• View Publication
• BIB
The relationship between the concepts of network and knowledge graph is explored. A knowledge graph can be considered a special type of network. When using a knowledge graph, various networks can be obtained from it, and network analysis procedures can be applied to them. RDF is a formalization of the knowledge graph concept for the Semantic Web, but some of its solutions are also extensible to a format for describing general networks.
Tutte's theorem as an educational formalization project
In this work, we present two results: The first result is the formalization of Tutte's theorem in Lean, a key theorem concerning matchings in graph theory. As this formalization is ready to be integrated in Lean's mathlib, it provides a valuable step in the path towards formalizing research-level mathematics in this area. The second result is a framework for doing educational formalization projects. This framework provides a structure to learn to formalize mathematics with minimal teacher input. This framework applies to both traditional academic settings and independent community-driven environments. We demonstrate the framework's use by connecting it to the process of formalizing Tutte's theorem.
On the Averaging Problem of Ideal Families Related to Frankl's Conjecture with Formal Proof by Lean 4
Frankl's conjecture, also known as the union-closed sets conjecture, can be equivalently expressed in terms of intersection-closed set families by considering the complements of sets. It posits that any family of sets closed under intersections, and containing both the ground set and the empty set, must have a ``rare vertex'' -- a vertex belonging to at most half of the members of the family. The concept of \emph{average rarity} describes a set family where the average degree of all the elements is at most half of the number of its members. Average rarity is a stronger property that implies the existence of a rare vertex. This paper focuses on ideal families, which are set families that are downward-closed (except the ground set) and include the ground set. We present a proof that the normalized degree sum of any ideal family is non-positive, which is equivalent to saying that every ideal family satisfies the average rarity condition. This proof is formalized and verified using the Lean 4 theorem prover.
Effective Computation of Generalized Abelian Complexity for Pisot Type Substitutive Sequences
Generalized abelian equivalence compares words by their factors up to a certain bounded length. The associated complexity function counts the equivalence classes for factors of a given size of an infinite sequence. How practical is this notion? When can these equivalence relations and complexity functions be computed efficiently? We study the fixed points of substitution of Pisot type. Each of their $k$-abelian complexities is bounded and the Parikh vectors of their length-$n$ prefixes form synchronized sequences in the associated Dumont--Thomas numeration system. Therefore, the $k$-abelian complexity of Pisot substitution fixed points is automatic in the same numeration system. Two effective generic construction approaches are investigated using the \texttt{Walnut} theorem prover and are applied to several examples. We obtain new properties of the Tribonacci sequence, such as a uniform bound for its factor balancedness together with a two-dimensional linear representation of its generalized abelian complexity functions.
Congruences on posets, relatively pseudocomplemented and Boolean posets
Published
• View Publication
• BIB
The aim of the present paper is to extend the concept of a congruence from lattices to posets. We use an approach different from that used by the first author and V. Snášel. By using our definition we show that congruence classes are convex. If the poset in question satisfies the Ascending Chain Condition as well as the Descending Chain Condition, then these classes turn out to be intervals. If the poset has a top element 1 then the 1-class of every congruence is a so-called strong filter. We study congruences on relatively pseudocomplemented posets which form a formalization of intuitionistic logic. For such posets we define so-called deductive systems and we show how they are connected with congruence kernels. We prove that every strong filter F of a relatively pseudocomplemented poset induces a congruence having F as its kernel. Finally, we consider Boolean posets which form a natural generalization of Boolean algebras. We show that congruences on Boolean posets in general do not share properties known from Boolean algebras, but congruence kernels of Boolean posets still have some interesting properties.
Using Walnut to solve problems from the OEIS
We use the automatic theorem prover Walnut to resolve various open problems from the OEIS and beyond. Specifically, we clarify the structure of sequence A260311, which concerns runs of sums of upper Wythoff numbers. We extend a result of Hajdu, Tijdeman, and Varga on polynomials with nonzero coefficients modulo a prime. Additionally, we settle open problems related to the anti-recurrence sequences A265389 and A299409, as well as the subsumfree sequences A026471 and A026475. Our findings also give rise to new open problems.
Generalized Hofstadter functions $G, H$ and beyond: numeration systems and discrepancy
Hofstadter's $G$ function is recursively defined via $G(0)=0$ and then $G(n)=n-G(G(n-1))$. Following Hofstadter, a family $(F_k)$ of similar functions is obtained by varying the number $k$ of nested recursive calls in this equation. We study here some Fibonacci-like sequences that are deeply connected with these functions $F_k$. In particular, the Zeckendorf theorem can be adapted to provide digital expansions via sums of terms of these sequences. On these digital expansions, the functions $F_k$ are acting as right shifts of the digits. These Fibonacci-like sequences can be expressed in terms of zeros of the polynomial $X^k{-}X^{k-1}{-}1$. Considering now the discrepancy of each function $F_k$, i.e., the maximal distance between $F_k$ and its linear equivalent, we retrieve the fact that this discrepancy is finite exactly when $k \le 4$. Thanks to that, we solve two twenty-year-old OEIS conjectures stating how close the functions $F_3$ and $F_4$ are from the integer parts of their linear equivalents. Moreover we establish that $F_k$ can coincide exactly with such an integer part only when $k\le 2$, while $F_k$ is almost additive exactly when $k \le 4$. Finally, a nice fractal shape a la Rauzy has been encountered when investigating the discrepancy of $F_3$. Almost all this article has been formalized and verified in the Coq/Rocq proof assistant.
The Critical Beta-splitting Random Tree III: The exchangeable partition representation and the fringe tree
In the critical beta-splitting model of a random $n$-leaf rooted tree, clades are recursively split into sub-clades, and a clade of $m$ leaves is split into sub-clades containing $i$ and $m-i$ leaves with probabilities $\propto 1/(i(m-i))$. Study of structure theory and explicit quantitative aspects of the model is an active research topic. It turns out that many results have several different proofs, and detailed studies of analytic proofs are given elsdewhere (via analysis of recursions and via Mellin transforms). This article describes two core probabilistic methods for studying $n \to \infty$ asymptotics of the basic finite-$n$-leaf models.
(i) There is a canonical embedding into a continuous-time model, that is a random tree CTCS(n) on $n$ leaves with real-valued edge lengths, and this model turns out to be more convenient to study. The family (CTCS(n), $n \ge 2)$ is consistent under a ``delete random leaf and prune" operation. That leads to an explicit inductive construction (the {\em growth algorithm}) of (CTCS(n), $n \ge 2)$ as $n$ increases, and then to a limit structure CTCS$(\infty)$ which can be formalized via exchangeable partitions, in some ways analogous to the Brownian continuum random tree.
(ii) There is an explicit description of the limit fringe distribution relative to a random leaf, whose graphical representation is essentially the format of the cladogram representation of biological phylogenies.
Machine Checked Proofs and Programs in Algebraic Combinatorics
Published in CPP 2025, Proceedings of the 14th ACM SIGPLAN International Conference on Certified Programs and Proofs
• View Publication
• BIB
We present a library of formalized results around symmetric functions and the character theory of symmetric groups. Written in Coq/Rocq and based on the Mathematical Components library, it covers a large part of the contents of a graduate level textbook in the field. The flagship result is a proof of the Littlewood-Richardson rule, which computes the structure constants of the algebra of symmetric function in the schur basis which are integer numbers appearing in various fields of mathematics, and which has a long history of wrong proofs. A specific feature of algebraic combinatorics is the constant interplay between algorithms and algebraic constructions: algorithms are not only in computations, but also are key ingredients in definitions and proofs. As such, the proof of the Littlewood-Richardson rule deeply relies on the understanding of the execution of the Robinson-Schensted algorithm. Many results in this library are effective and actually used in computer algebra systems, and we discuss their certified implementation.
Simulating NMR Spectra with a Quantum Computer
The procedure for simulating the nuclear magnetic resonance spectrum linked to the spin system of a molecule for a certain nucleus entails diagonalizing the associated Hamiltonian matrix. As the dimensions of said matrix grow exponentially with respect to the spin system's atom count, the calculation of the eigenvalues and eigenvectors marks the performance of the overall process. The aim of this paper is to provide a formalization of the complete procedure of the simulation of a spin system's NMR spectrum while also explaining how to diagonalize the Hamiltonian matrix with a quantum computer, thus enhancing the overall process's performance. Two well-known quantum algorithms for calculating the eigenvalues of a matrix are analyzed and put to the test in this context: quantum phase estimation and the variational quantum eigensolver. Additionally, we present simulated results for the later approach while also addressing the hypothetical noise found in a physical quantum computer.
The Repetition Threshold for Rote Sequences
We consider Rote words, which are infinite binary words with factor complexity $2n$. We prove that the repetition threshold for this class is $5/2$. Our technique is purely computational, using the Walnut theorem prover and a new technique for generating automata from morphisms due to the first author and his co-authors.
On de Bruijn Rings and Families of Almost Perfect Maps
Published
• View Publication
• BIB
De Bruijn tori, or perfect maps, are two-dimensional periodic arrays of letters from a finite alphabet, where each possible pattern of shape (m,n) appears exactly once in a single period. While the existence of certain de Bruijn tori, such as square tori with odd m=n element {3,5,7} and even alphabet sizes, remains unresolved, sub-perfect maps are often sufficient in applications like positional coding. These maps capture a large number of patterns, with each appearing at most once. While previous methods for generating such sub-perfect maps cover only a fraction of the possible patterns, we present a construction method for generating almost perfect maps for arbitrary pattern shapes and arbitrary non-prime alphabet sizes, including the above mentioned square tori with odd m=n element {3,5,7} as long that the alphabet size is non-prime. This is achieved through the introduction of de Bruijn rings, a minimal-height sub-perfect map and a formalization of the concept of families of almost perfect maps. The generated sub-perfect maps are easily decodable which makes them perfectly suitable for positional coding applications.
Homeostasis in Input-Output Networks: Structure, Classification and Applications
Published
• View Publication
• BIB
Homeostasis is concerned with regulatory mechanisms, present in biological systems, where some specific variable is kept close to a set value as some external disturbance affects the system. Mathematically, the notion of homeostasis can be formalized in terms of an input-output function that maps the parameter representing the external disturbance to the output variable that must be kept within a fairly narrow range. This observation inspired the introduction of the notion of infinitesimal homeostasis, namely, the derivative of the input-output function is zero at an isolated point. This point of view allows for the application of methods from singularity theory to characterize infinitesimal homeostasis points (i.e. critical points of the input-output function). In this paper we review the infinitesimal approach to the study of homeostasis in input-output networks. An input-output network is a network with two distinguished nodes `input' and `output', and the dynamics of the network determines the corresponding input-output function of the system. This class of dynamical systems provides an appropriate framework to study homeostasis and several important biological systems can be formulated in this context. Moreover, this approach, coupled to graph-theoretic ideas from combinatorial matrix theory, provides a systematic way for classifying different types of homeostasis (homeostatic mechanisms) in input-output networks, in terms of the network topology. In turn, this leads to new mathematical concepts, such as, homeostasis subnetworks, homeostasis patterns, homeostasis mode interaction. We illustrate the usefulness of this theory with several biological examples: biochemical networks, chemical reaction networks (CRN), gene regulatory networks (GRN), Intracellular metal ion regulation and so on.
Edge-length preserving embeddings of graphs between normed spaces
Published
• View Publication
• BIB
The concept of graph flattenability, initially formalized by Belk and Connelly and later expanded by Sitharam and Willoughby, extends the question of embedding finite metric spaces into a given normed space. A finite simple graph $G=(V,E)$ is said to be $(X,Y)$-flattenable if any set of induced edge lengths from an embedding of $G$ into a normed space $Y$ can also be realised by an embedding of $G$ into a normed space $X$. This property, being minor-closed, can be characterized by a finite list of forbidden minors. Following the establishment of fundamental results about $(X,Y)$-flattenability, we identify sufficient conditions under which it implies independence with respect to the associated rigidity matroids for $X$ and $Y$. We show that the spaces $\ell_2$ and $\ell_\infty$ serve as two natural extreme spaces of flattenability and discuss $(X, \ell_p )$-flattenability for varying $p$. We provide a complete characterization of $(X,Y)$-flattenable graphs for the specific case when $X$ is 2-dimensional and $Y$ is infinite-dimensional.
A Formal Proof of R(4,5)=25
In 1995, McKay and Radziszowski proved that the Ramsey number R(4,5) is equal to 25. Their proof relies on a combination of high-level arguments and computational steps. The authors have performed the computational parts of the proof with different implementations in order to reduce the possibility of an error in their programs. In this work, we prove this theorem in the interactive theorem prover HOL4 limiting the uncertainty to the small HOL4 kernel. Instead of verifying their algorithms directly, we rely on the HOL4 interface to MiniSat SAT to prove gluing lemmas. To reduce the number of such lemmas and thus make the computational part of the proof feasible, we implement a generalization algorithm. We verify that its output covers all the possible cases by implementing a custom SAT-solver extended with a graph isomorphism checker.
Exploring the Crochemore and Ziv-Lempel factorizations of some automatic sequences with the software Walnut
We explore the Ziv-Lempel and Crochemore factorizations of some classical automatic sequences making an extensive use of the theorem prover Walnut.
Complexity and asymptotics of structure constants
Published
• View Publication
• BIB
Kostka, Littlewood-Richardson, Kronecker, and plethysm coefficients are fundamental quantities in algebraic combinatorics, yet many natural questions about them stay unanswered for more than 80 years. Kronecker and plethysm coefficients lack ``nice formulas'', a notion that can be formalized using computational complexity theory. Beyond formulas and combinatorial interpretations, we can attempt to understand their asymptotic behavior in various regimes, and inequalities they could satisfy. Understanding these quantities has applications beyond combinatorics. On the one hand, the asymptotics of structure constants is closely related to understanding the [limit] behavior of vertex and tiling models in statistical mechanics. More recently, these structure constants have been involved in establishing computational complexity lower bounds and separation of complexity classes like VP vs VNP, the algebraic analogs of P vs NP in arithmetic complexity theory. Here we discuss the outstanding problems related to asymptotics, positivity, and complexity of structure constants focusing mostly on the Kronecker coefficients of the symmetric group and, less so, on the plethysm coefficients.
This expository paper is based on the talk presented at the Open Problems in Algebraic Combinatorics coneference in May 2022.
Signal processing on large networks with group symmetries
Current methods of graph signal processing rely heavily on the specific structure of the underlying network: the shift operator and the graph Fourier transform are both derived directly from a specific graph. In many cases, the network is subject to error or natural changes over time. This motivated a new perspective on GSP, where the signal processing framework is developed for an entire class of graphs with similar structures. This approach can be formalized via the theory of graph limits, where graphs are considered as random samples from a distribution represented by a graphon.
When the network under consideration has underlying symmetries, they may be modeled as samples from Cayley graphons. In Cayley graphons, vertices are sampled from a group, and the link probability between two vertices is determined by a function of the two corresponding group elements. Infinite groups such as the 1-dimensional torus can be used to model networks with an underlying spatial reality. Cayley graphons on finite groups give rise to a Stochastic Block Model, where the link probabilities between blocks form a (edge-weighted) Cayley graph. This manuscript summarizes some work on graph signal processing on large networks, in particular samples of Cayley graphons.