Score
Creates precise mathematical definitions and formal specifications for concepts by translating informal or intuitive descriptions into formal language. This includes generalizing or refining existing definitions and proving consistency, equivalence, or other formal properties.
This work addresses the automatic formalization of informal mathematical statements into proof languages (e.g., Lean, Isabelle), aiming to enhance theorem proving automation and improve the verifiability of large language model (LLM) mathematical reasoning. Methodologically, it integrates formal logic, mathematical knowledge representation, LLM fine-tuning and prompt engineering, cross-modal alignment, and formal verification toolchains into an end-to-end automated formalization framework. Its contributions are threefold: first, it establishes automated formalization as a novel paradigm for enhancing the trustworthiness of LLM mathematical reasoning, unifying technical developments from both mathematical logic and LLM perspectives; second, it systematically categorizes open-source models, benchmark datasets, and core challenges, presenting the most comprehensive research landscape to date; third, it identifies critical technical bottlenecks and outlines concrete development pathways, thereby fostering synergistic advancement in automated theorem proving and trustworthy LLM-based mathematical reasoning.
This work addresses the poor readability of formal proofs, which hinders comprehension by mathematicians. We propose a structure-aware recursive summarization framework that leverages large language models to generate stepwise, informal, and hierarchical natural-language summaries of formal proofs—e.g., those written in Lean. The method recursively compresses subproofs along the proof dependency graph and then integrates contextual information to produce coherent, natural-language explanations. Crucially, it achieves end-to-end generation of highly readable natural-language proofs while strictly preserving logical fidelity. Experiments on textbook-level theorems and the Lean Mathematical Library demonstrate that the generated summaries match or surpass human-written reference proofs in readability, logical faithfulness, and mathematical rigor. These results validate both the method’s effectiveness and its generalizability across diverse mathematical domains.
This work investigates the capability of large language models (LLMs) to automatically formalize real-world mathematical definitions—sourced from Wikipedia and arXiv—into Isabelle/HOL. To address the limitation of existing benchmarks (e.g., miniF2F) in capturing definition-specific challenges, we introduce Def_Wiki/Def_ArXiv, the first dual-source benchmark for mathematical definition formalization. We propose definition grounding—a method that integrates context-aware prompting with proof-assistant feedback-driven structured refinement—and close the loop via external verification. Experiments demonstrate a 16% improvement in model self-correction ability and a 43% reduction in undefined-symbol errors. Our findings reveal that formalizing mathematical definitions is substantially more challenging than theorem proving, exposing critical bottlenecks in LLMs’ semantic precision and symbolic consistency.
This work addresses the challenge of reliably translating formalized mathematics into natural language that is both precise and readable. It proposes a symbol-to-text framework based on an intermediate language architecture, extending conventional syntactic sugar mechanisms to support general mathematical expressions. By unifying formal systems such as Agda, Lean, and Rocq through Dedukti and integrating Grammatical Framework (GF), the approach ensures grammatical correctness and expressive diversity across multiple languages. The resulting system, Informath, generates fluent, accurate, and multilingual mathematical narratives at relatively low development cost, effectively rendering AI-generated or automatically formalized proofs into comprehensible expository text.
Current formal reasoning systems face a critical bottleneck: the scarcity of aligned parallel corpora between natural language (NL) and formal proof languages such as Lean. This work introduces Herald, the first large-scale, high-quality NL–Lean 4 aligned dataset, covering core content from Mathlib4 and structured via a hierarchical, retrievable section-level alignment framework. We propose a dual-path enhancement pipeline—tactic-driven and informal-description–guided—built upon the Lean-jixia analyzer. To our knowledge, this is the first system enabling fully automated formalization of graduate-level mathematics textbook content. Fine-tuning on Herald yields the Herald Translator model, which achieves 93.2% Pass@128 on miniF2F-test—substantially outperforming prior baselines—and successfully formalizes the template chapter of the Stack Project. Both the Herald dataset and the Herald Translator model are publicly released.
Formal program specifications are notoriously difficult, error-prone, and inefficient to write manually. To address this, we propose a two-stage LLM-driven approach: dialogue-guided specification synthesis followed by mutation-based verification. First, multi-turn dialogues model complex semantic requirements; second, four mutation operators—insertion, replacement, deletion, and reordering—enable verifiability-driven selection, eliminating reliance on rigid templates or syntactic grammars. Our method integrates code understanding, prompt engineering, and heuristic verifiability assessment. Evaluated on SV-COMP and a custom Java benchmark comprising 385 programs, it generates 279 verifiable specifications. These achieve significantly higher completeness and accuracy than pure-LLM baselines and classical tools (e.g., Houdini, Daikon). To our knowledge, this is the first approach to achieve both high coverage and formal verifiability in fully automated specification generation.
Existing approaches to automatic formalization struggle to scale to entire mathematical textbooks due to challenges such as cross-file dependencies, import resolution, and end-to-end compilation. This work proposes the M2F framework, the first system capable of project-scale automatic formalization of mathematical literature. M2F operates in two stages: first, it constructs compilable theorem skeletons by performing dependency-aware ordering and declaration repair; second, it completes proofs through goal-conditioned local editing, iteratively refined via closed-loop feedback from the Lean proof checker. Applied to a 479-page textbook on real and convex analysis, M2F generated 153,853 lines of Lean code within three weeks, achieving a proof completion rate of 96%—substantially outperforming both the 80% baseline and manual formalization efforts in efficiency.
This work addresses inconsistencies arising from structural mismatches between natural language and formal languages during requirements formalization. It proposes a “consistency through formalization” principle, mandating strict logical alignment among natural language, the structured language FRETish, and the formal temporal logic MTL. Guided by this principle, the authors refine the FRETish-to-MTL translation pipeline in NASA’s FRET tool. Their approach uniquely integrates cross-layer consistency constraints into a collaborative framework combining large language models and formal verification tools. This integration not only uncovers and corrects multiple inconsistencies in the original translation but also demonstrates superior correctness and reliability, as substantiated by formal equivalence proofs and empirical evaluation.
This work addresses the challenge of achieving reliable, low-cost automatic formalization of mathematical proofs under limited computational resources. It introduces Trellis, a novel system that translates mathematicians’ intuitive notion of “rigor” into executable process semantics by constructing a deterministic, constraint-guided workflow based on general-purpose large language model agents. Without requiring domain-specific training, Trellis employs an iterative refinement mechanism to progressively transform informal natural language proofs into formal Lean proofs. The system demonstrates its efficacy and practicality by successfully formalizing, in an end-to-end manner, a recent breakthrough result in Ramsey theory, thereby validating its capacity to bridge informal mathematical reasoning and machine-checkable formalization.
Existing approaches to automatic formalization are largely confined to isolated statements and struggle to capture the intricate dependency structures among axioms, definitions, and lemmas within mathematical theories. This work introduces a novel paradigm—*theory-level automatic formalization*—which systematically advocates shifting from statement-level to integrated theory-level formalization. By constructing formalized mathematical libraries, modeling dependency graphs, and establishing mappings from natural to formal languages, the approach enables machine-verifiable translation of entire mathematical theories along with their internal structures. The paper delineates core challenges inherent to this direction, proposes three viable research pathways, and releases a comprehensive survey resource to catalyze further progress in the field.