arXiv ScienceSearch

arXiv · 2601.03298

130k Lines of Formal Topology in Two Weeks: Simple and Cheap Autoformalization for Everyone?

Abstract

This is a brief description of a project that has already autoformalized a large portion of the general topology from the Munkres textbook (which has in total 241 pages in 7 chapters and 39 sections). The project has been running since November 21, 2025 and has as of January 4, 2026, produced 160k lines of formalized topology. Most of it (about 130k lines) have been done in two weeks,from December 22 to January 4, for an LLM subscription cost of about \$100. This includes a 3k-line proof of Urysohn's lemma, a 2k-line proof of Urysohn's Metrization theorem, over 10k-line proof of the Tietze extension theorem, and many more (in total over 1.5k lemmas/theorems). The approach is quite simple and cheap: build a long-running feedback loop between an LLM and a reasonably fast proof checker equipped with a core foundational library. The LLM is now instantiated as ChatGPT (mostly 5.2) or Claude Sonnet (4.5) run through the respective Codex or Claude Code command line interfaces. The proof checker is Chad Brown's higher-order set theory system Megalodon, and the core library is Brown's formalization of basic set theory and surreal numbers (including reals, etc). The rest is some prompt engineering and technical choices which we describe here. Based on the fast progress, low cost, virtually unknown ITP/library, and the simple setup available to everyone, we believe that (auto)formalization may become quite easy and ubiquitous in 2026, regardless of which proof assistant is used.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Josef Urban. 2026-01-06. 130k Lines of Formal Topology in Two Weeks: Simple and Cheap Autoformalization for Everyone?. https://arxiv.org/abs/2601.03298

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Differential Equations as Fixpoints and Games

Games and fixpoints are unified by proving that first-order game logic GL and the first-order modal mu-calculus L_mu are proved to be equiexpressive and equivalent, thereby fully aligning their expressive and deductive power. That is, there is a semantics-preserving translation from GL to L_mu, and vice versa. And both translations are provability-preserving, while equivalence with there-and-back-again roundtrip translations are provable in both calculi. This is to be contrasted with the propositional case, where game logic is strictly less expressive than the modal mu-calculus (without adding sabotage games). The extensions with differential equations, differential game logic (dGL) and differential modal mu-calculus, are also proved equiexpressive and equivalent. Moreover, as the continuous dynamics are definable by fixpoints or via games, ODEs can be axiomatized completely and, as a consequence, infinitesimally robust properties of ODEs can be decided via proof search. Rational gameplay provably collapses the games into single-player games to yield a strong arithmetical completeness theorem for dGL with rational-time ODEs.

cs.LO

Self-extensional logics of formal inconsistency: Decidability and limits for paraconsistency

RmbC is a self-extensional paraconsistent logic in the family of Logics of Formal Inconsistency (LFIs). This system is obtained from mbC (the basic LFI) by adding the replacement property via two global inference rules. RmbC is characterized by a non-explosive negation $\neg$ and a consistency operator $\circ$, which recovers the principle of explosion in a controlled way. Together with its principal axiomatic extensions, RmbC admits a standard Lindenbaum-Tarski algebraization, with Boolean algebras with LFI operators (BALFIs) as its algebraic semantics. In this paper, we study how far this self-extensional paraconsistent behavior can be extended axiomatically, starting from RmbC. We classify pairs of very natural consistency axioms according to whether they preserve paraconsistency or force classical collapse; identify six minimal explosive combinations that collapse to a single algebraic core; and isolate a separate structural obstruction for the combination of excluded middle for $\neg$ with an involutive negation. We also investigate, for the first time, the decidability of this family of self-extensional LFIs. As a first result, we prove the finite model property for RmbC with respect to BALFI semantics via an algebraic filtration, which yields decidability, and transfer this result to several paraconsistent axiomatic extensions of RmbC. Finally, we establish a coNEXPTIME upper bound for the validity problem of RmbC and a coNP-hardness lower bound, and prove coNP-completeness for the principal extensions containing one of the six minimal explosive pairs.

cs.LO

The Stochastic Target Discounted-Sum Problem

The target discounted-sum problem (TDS) asks, given a finite integer alphabet $Σ$, a rational discount factor $λ$, and a rational target $t$, whether some infinite sequence over $Σ$ has discounted sum exactly $t$. This problem remains open and underlies several open questions in automata theory, games, and Markov decision processes. We introduce and solve its stochastic counterpart, the stochastic target discounted-sum problem, which replaces existence by computation of the probability. We show that the probability that a random sequence generated by a finite Markov chain has discounted sum $t$ is rational and computable in pseudo-polynomial time. We further show how to decide, in polynomial time, whether the discounted-sum distribution of a Markov chain is atomless, and how to approximate to an arbitrary precision the probability that the discounted sum exceeds a rational threshold. Our techniques for the stochastic TDS problem allow us to make progress on TDS objectives in stochastic games, which are known to be as hard as the TDS problem. Restricting the maximizing player to finite-memory strategies, while allowing the minimizing player to use arbitrary strategies, we reduce the value problem and the synthesis problem to corresponding problems for safety objectives in stochastic games. This yields computable optimal values and deterministic optimal strategies with pseudo-polynomially bounded memory for stochastic games, and results in pseudo-polynomial-time algorithms for special cases of Markov decision processes and deterministic two-player games.

cs.LO