🤖 AI Summary
This study addresses the absence of a unified task definition and systematic survey in the automatic formalization of first-order logic (FOL). It establishes a principled definition of FOL autoformalization by explicitly decoupling the ontology extraction and logical translation phases for the first time. Building upon this framework, the paper systematically reviews existing datasets, evaluation metrics, and large language model-based approaches—including fine-tuning, prompt engineering, and verification-guided refinement—for natural language-to-FOL translation. Furthermore, this work clarifies inconsistencies in cross-study evaluations and bridges critical theoretical gaps in the field. By identifying open challenges in benchmarking, semantic evaluation, and end-to-end applications, it provides a clear roadmap to guide future research in automated FOL formalization.
📝 Abstract
Large Language Models (LLMs) have renewed interest in autoformalization. Yet, when First-Order Logic (FOL) is considered as the target formalism, the field still lacks a unified task formulation and a systematic survey. This paper addresses this gap: we first provide a principled definition for the FOL-autoformalization task by distinguishing Ontology Extraction from Logical Translation, showing how their conflation obscures (cross-study) evaluation; we review existing datasets, evaluation metrics, and LLM-based methods, including fine-tuning, prompting, and verification-based refinement; we identify open challenges in benchmarking, semantic evaluation, ontology-aware methods, and end-to-end applications.