Herald: A Natural Language Annotated Lean 4 Dataset

📅 2024-10-09
🏛️ arXiv.org
📈 Citations: 3
✨ Influential: 1
📄 PDF
🤖 AI Summary
Current formal reasoning systems face a critical bottleneck: the scarcity of aligned parallel corpora between natural language (NL) and formal proof languages such as Lean. This work introduces Herald, the first large-scale, high-quality NL–Lean 4 aligned dataset, covering core content from Mathlib4 and structured via a hierarchical, retrievable section-level alignment framework. We propose a dual-path enhancement pipeline—tactic-driven and informal-description–guided—built upon the Lean-jixia analyzer. To our knowledge, this is the first system enabling fully automated formalization of graduate-level mathematics textbook content. Fine-tuning on Herald yields the Herald Translator model, which achieves 93.2% Pass@128 on miniF2F-test—substantially outperforming prior baselines—and successfully formalizes the template chapter of the Stack Project. Both the Herald dataset and the Herald Translator model are publicly released.

Technology Category

Knowledge Representation and Reasoning: Automated Reasoning and Theorem ProvingNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.Machine Learning: Large Multimodal Models (LMMs)

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Verifiable formal languages like Lean have profoundly impacted mathematical reasoning, particularly through the use of large language models (LLMs) for automated reasoning. A significant challenge in training LLMs for these formal languages is the lack of parallel datasets that align natural language with formal language proofs. To address this challenge, this paper introduces a novel framework for translating the Mathlib4 corpus (a unified library of mathematics in formal language Lean 4) into natural language. Building upon this, we employ a dual augmentation strategy that combines tactic-based and informal-based approaches, leveraging the Lean-jixia system, a Lean 4 analyzer. We present the results of this pipeline on Mathlib4 as Herald (Hierarchy and Retrieval-based Translated Lean Dataset). We also propose the Herald Translator, which is fine-tuned on Herald. Herald translator achieves a 93.2% accuracy (Pass@128) on formalizing statements in the miniF2F-test and a 22.5% accuracy on our internal graduate-level textbook dataset, outperforming InternLM2-Math-Plus-7B (74.0% and 7.5%) and TheoremLlama (50.1% and 4.0%). Furthermore, we propose a section-level translation framework for real-world applications. As a direct application of Herald translator, we have successfully translated a template section in the Stack project, marking a notable progress in the automatic formalization of graduate-level mathematical literature. Our model, along with the datasets, are open-sourced to the public.
Problem

Research questions and friction points this paper is trying to address.

Lack of parallel datasets for Lean 4
Translation of Mathlib4 into natural language
Automated formalization of graduate-level mathematics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Translating Mathlib4 to natural language
Dual augmentation strategy with Lean-jixia
Herald Translator for formalization accuracy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Peking University | National University of Singapore
G
Guoxiong Gao
Peking University
Y
Yutong Wang
National University of Singapore
J
Jiedong Jiang
Peking University
Q
Qi Gao
Peking University
Z
Zihan Qin
Peking University
Tianyi Xu
Tianyi Xu
Tulane University
Reinforcement LearningNetwork OptimizaitonStatisticsNLP(LLM)Operations research
B
Bin Dong
Peking University