Score
Designs, builds, and analyzes formal mathematical artifacts such as proofs, models, and quantitative relationships; formulates and verifies rigorous arguments, derives formulas and transformations, constructs algebraic, calculus-based, discrete, or statistical models, and develops analytical methods for solving equations, optimizing functions, and producing symbolic or numeric results.
Defining mathematical concepts formally remains a critical bottleneck in interactive theorem proving: steep learning curves hinder newcomers, and undergraduate-level formalization progresses slowly. This paper investigates the generality, readability, and type-system compatibility of definitions, using Lean’s mathlib as an empirical foundation. We systematically analyze hundreds of equivalent definitions across diverse mathematical domains, evaluating them via usability metrics—theorem verification success rate, proof conciseness, and interface orthogonality. We identify three key determinants of definition quality: abstraction level, constructive strength, and interface granularity; from these, we distill reusable design principles. Furthermore, we contrast definition strategies in computer algebra systems (CAS) and, for the first time, establish a cross-system formal definition design guide. Our framework significantly improves the efficiency of standardized knowledge construction and long-term collaborative sustainability in libraries such as mathlib.
Large language models (LLMs) face challenges in formal mathematical verification—including capability coupling, coarse-grained evaluation, and scarcity of high-quality, language-diverse training data. Method: We systematically decouple formal verification into six fine-grained subtasks (e.g., specification translation, proof completion) and construct FM-alpaca, a 18K-sample high-quality instruction-response dataset covering five mainstream formal languages: Coq, Lean4, Dafny, ACSL, and TLA+. Leveraging GPT-4o distillation and supervised fine-tuning (SFT), we propose FM-Bench—the first cross-language, task-decoupled benchmark for formal verification. Contribution/Results: Empirical results show that fine-tuning on formalization data significantly improves formal verification performance (up to 2.9× gain) and positively transfers to mathematical reasoning and programming tasks. Both the model and benchmark are publicly released.
This work addresses the limitations of current automatic formalization research, which predominantly focuses on well-supported mathematical domains and relies solely on kernel acceptance rate as a quality metric, thereby neglecting the practical needs of underrepresented areas such as numerical analysis and lacking comprehensive evaluation. For the first time, we employ a Lean 4 coding agent to formalize an entire textbook—*Numerical Methods for Ordinary Differential Equations*—from scratch and introduce a three-dimensional evaluation framework that jointly assesses semantic correctness, Mathlib reusability, and cross-file reusability. Through LLM-as-judge, semantic validation, and dependency analysis, we uncover pervasive issues in existing systems, including incomplete statements and weakened assumptions, demonstrating that kernel acceptance rate substantially overestimates formalization quality. Our approach establishes a reproducible, multidimensional auditing paradigm for trustworthy automated formalization.
This study presents the first systematic evaluation of state-of-the-art large language models’ ability to generate formally verifiable proofs in graduate-level mathematics, including analysis, algebra, probability, and logic. The authors construct a private benchmark that pairs natural language mathematical problems with their formalized statements in Lean 4 and employ an agent-based framework to automatically assess whether model-generated proofs are accepted by the Lean 4 proof checker. Experimental results show that the best base model achieves an accuracy of 33.5%, while other models exhibit substantially lower performance. The work further provides a detailed empirical analysis of tool usage, failure modes, reasoning costs, and latency, offering valuable insights into the current capabilities and limitations of large language models in formal mathematical reasoning.
This study investigates the algebraicity and arithmetic properties of hypergeometric functions over the rational numbers, finite fields, and p-adic fields. Leveraging the SageMath computer algebra system, the work integrates techniques from algebraic number theory, finite field theory, and p-adic analysis to systematically implement, for the first time in an open-source framework, algorithms capable of determining algebraicity, computing valuations, and solving for minimal polynomials in positive characteristic. This implementation fills a critical gap in existing computational toolchains by enabling uniform arithmetic analysis of hypergeometric functions across multiple number-theoretic domains, thereby substantially enhancing SageMath’s capacity for algebraic manipulation of such functions.
This work addresses the challenges of data scarcity and the difficulty of simultaneously satisfying layout constraints and geometric precision in multimodal analytic geometry problems. We propose the first neuro-symbolic framework that bridges free-form text and precise signed distance field (SDF)-based diagram generation through a formal intermediate language, Coordinate Description Language (CDL). A closed-loop pipeline orchestrated by four large language models enables fully automated problem generation, formalization, measurement, and verification—without requiring human annotations—significantly enhancing the accuracy and scalability of geometric representations. Using this approach, we construct AnalyticGeo7K, a dataset comprising over 7,000 samples, achieving a median relative error of only 0.70% and ensuring that 82.3% of answers exhibit errors within 5%.
Current automated formalization tools struggle to independently handle the formalization of complex mathematical proofs. This study investigates how human experts conduct proof formalization with AI assistance through a mixed-methods approach, combining qualitative inquiry with controlled user experiments across diverse domains and difficulty levels. It provides the first systematic characterization of how users flexibly orchestrate multiple AI tools in real-world scenarios and reveals a central human need to retain high-level control in human-AI collaboration. The findings demonstrate that AI assistance significantly improves formalization accuracy, and users consistently adapt their tool usage dynamically based on task requirements.