Score
Designs and writes precise mathematical problem and model specifications by translating informal goals into formal objective functions, constraints, and hypothesis statements. Builds and analyzes mathematical formulations for optimization, estimation, or inference—specifying objective criteria, feasible sets, and the mathematical structure needed for algorithmic solution or statistical testing.
Mathematical optimization modeling heavily relies on domain experts and suffers from low automation. Method: This work systematically investigates how large language models (LLMs) can empower automated mathematical modeling, focusing on data synthesis, instruction fine-tuning, reasoning framework design, benchmark construction, and evaluation methodology. To address pervasive labeling errors (>40%) in mainstream benchmarks (e.g., OptiMath, MOBench), we conduct the first large-scale manual verification and cleaning, yielding the high-quality OptiClean dataset. Contribution/Results: Based on OptiClean, we establish the first fair, reproducible automated modeling leaderboard; release an open-source repository integrating datasets, code, literature, and an online evaluation platform; and provide a standardized evaluation framework, reliable benchmark, and scalable technical paradigm for LLM-driven modeling automation—significantly advancing the field’s standardization and rigor.
To address the accuracy bottleneck in automatic natural language-to-mixed-integer linear programming (NL-to-MILP) modeling—stemming from scarce high-quality annotated data and insufficient integration of domain expertise—this paper proposes an optimization-knowledge-enhanced large language model (LLM) framework. Our method comprises three core components: (1) a fine-grained, category-specific error analysis–driven data cleaning strategy; (2) a MILP-semantic-structured, class-aware multi-turn reasoning prompting framework; and (3) an iterative validation and refinement mechanism incorporating solver feedback. Extensive experiments across multiple foundational LLMs demonstrate an average 14.2-percentage-point improvement in modeling accuracy. Notably, robustness is significantly enhanced on critical subtasks—including complex constraint formulation and integer variable identification. The proposed approach establishes a new, interpretable, and solver-verified paradigm for AI-driven operations research modeling.
This work addresses the autoformulation problem—automatically translating natural-language problem descriptions into solvable mathematical optimization models. We propose the first LLM-driven Monte Carlo Tree Search (MCTS) framework for this task, enabling dynamic hypothesis generation and formal correctness evaluation. Our method integrates hierarchical optimization modeling representations, LLM-based semantic understanding, and MCTS-based search strategies. A key innovation is an equivalence-aware pruning mechanism that reduces search overhead by over 40%. Empirically, our approach achieves state-of-the-art performance on LP/MIP benchmarks, outperforming all existing baselines. LLM-assisted verification accelerates correctness assessment significantly. Moreover, this work formally defines the autoformulation task for the first time, establishing a scalable, automated paradigm to lower the barrier to optimization modeling and empower domain experts.
Large language models (LLMs) face challenges in formal mathematical verification—including capability coupling, coarse-grained evaluation, and scarcity of high-quality, language-diverse training data. Method: We systematically decouple formal verification into six fine-grained subtasks (e.g., specification translation, proof completion) and construct FM-alpaca, a 18K-sample high-quality instruction-response dataset covering five mainstream formal languages: Coq, Lean4, Dafny, ACSL, and TLA+. Leveraging GPT-4o distillation and supervised fine-tuning (SFT), we propose FM-Bench—the first cross-language, task-decoupled benchmark for formal verification. Contribution/Results: Empirical results show that fine-tuning on formalization data significantly improves formal verification performance (up to 2.9× gain) and positively transfers to mathematical reasoning and programming tasks. Both the model and benchmark are publicly released.
Automating the formalization of research-level mathematical theorems in the Lean proof assistant remains challenging due to the gap between abstract mathematical structures and their concrete instantiations. Method: We propose a structured, template-driven approach that bridges this gap systematically. It employs reusable, modular templates to explicitly encode mappings from abstract structures to concrete instances; leverages large language models to generate candidate definitions and theorems; utilizes Lean’s type-class mechanism for automatic instance resolution; and incorporates structural hypothesis verification and feedback-guided iterative refinement to ensure formal correctness. Contribution/Results: This work achieves the first end-to-end automated formalization of theorems across multiple concrete instances derived from a single abstract structure. Evaluated on an optimization-theory dataset, our method successfully generated multiple correct, machine-verifiable Lean proofs. It significantly improves both the efficiency and breadth of mathematical formalization, advancing scalable, reliable automation in interactive theorem proving.
Existing machine learning frameworks suffer from insufficient formalization of objective functions and lack a unified, cross-domain behavioral design paradigm. Method: We propose an equation-constrained compositional function modeling approach for learners, constructing task graphs and compositional semantic graphs to enable model-agnostic behavioral specification and optimization. We introduce a novel task-oriented pattern language framework and the “manipulator” task paradigm, supporting end-to-end, architecture-agnostic, and adversarial-training-free minimal editing of data attributes. Contribution/Results: Theoretically, our work integrates formal methods and theoretical computer science principles. Empirically, we demonstrate precise, controllable, and interpretable behavioral editing on small-scale models under stable training—without stochastic sampling or data intervention—yielding significant improvements in deployment efficiency and formal verifiability.
Existing approaches to automatic formalization struggle to scale to entire mathematical textbooks due to challenges such as cross-file dependencies, import resolution, and end-to-end compilation. This work proposes the M2F framework, the first system capable of project-scale automatic formalization of mathematical literature. M2F operates in two stages: first, it constructs compilable theorem skeletons by performing dependency-aware ordering and declaration repair; second, it completes proofs through goal-conditioned local editing, iteratively refined via closed-loop feedback from the Lean proof checker. Applied to a 479-page textbook on real and convex analysis, M2F generated 153,853 lines of Lean code within three weeks, achieving a proof completion rate of 96%—substantially outperforming both the 80% baseline and manual formalization efforts in efficiency.
This work addresses the challenges posed by higher-order functions in mathematical optimization modeling, which often lead to unnatural LaTeX output and inefficient constraint verification. To overcome these issues, the authors propose an egglog-based optimization approach that performs desugaring reconstruction on the λ-calculus intermediate representation of JijModeling 2, thereby recovering comprehension-like syntactic structures. By integrating Henkin-style constants with Datalog-inspired rules, the method enables declarative, multi-step constraint checking. A custom cost model and equality saturation techniques are introduced to enhance performance significantly: the generated LaTeX aligns more closely with conventional mathematical notation, and complex constraint validation—previously requiring minutes or failing to terminate—now completes within seconds.
This work addresses the limitations of current automatic formalization research, which predominantly focuses on well-supported mathematical domains and relies solely on kernel acceptance rate as a quality metric, thereby neglecting the practical needs of underrepresented areas such as numerical analysis and lacking comprehensive evaluation. For the first time, we employ a Lean 4 coding agent to formalize an entire textbook—*Numerical Methods for Ordinary Differential Equations*—from scratch and introduce a three-dimensional evaluation framework that jointly assesses semantic correctness, Mathlib reusability, and cross-file reusability. Through LLM-as-judge, semantic validation, and dependency analysis, we uncover pervasive issues in existing systems, including incomplete statements and weakened assumptions, demonstrating that kernel acceptance rate substantially overestimates formalization quality. Our approach establishes a reproducible, multidimensional auditing paradigm for trustworthy automated formalization.