🤖 AI Summary
This study addresses the limitation of traditional autoformalization methods that rely on fine-grained feedback, which often triggers cross-level verification conflicts. To this end, we propose ProGS, a novel approach that introduces a tree-structured proof sketch to guide the construction and repair of formal models using large language models. By precisely mapping verification failures to specific nodes within the proof tree, ProGS enables structured iterative refinement while synergizing with formal tools to enhance automation. Experimental evaluations across 27 benchmark systems demonstrate that our method significantly improves syntactic validity, deductive verifiability, and behavioral correctness compared to existing baselines.
📝 Abstract
Formal modeling provides strong guarantees about system correctness, but developing and repairing formal models remains labor-intensive and requires substantial expertise in logic and formal reasoning. Recent LLM-based autoformalization agents seek to reduce this burden by generating candidate formal models and revising them using feedback from formal tools. However, the existing approaches follow a generate-and-repair paradigm, in which repairs are driven by verification failures of the generated model and therefore depend heavily on both the granularity of the feedback and the LLM's repair capability. As a consequence, a repair targeting one level of verification may invalidate properties at another level, which requires reasoning over the complete set of event guards. To address these limitations, we propose Proof-Sketch-Guided Formal Model Synthesis (ProGS), an autoformalization method centered on model-based proof sketches. A model-based proof sketch represents the proof structure of the target formal system as a tree. Internal nodes capture case splits and inductive reasoning steps, while leaf nodes correspond to concrete state-transition events that realize individual subgoals. ProGS uses LLMs to generate and repair these sketches, with verification failures mapped back to specific nodes and subtrees to provide structured guidance for iterative repair. Our evaluation on a benchmark of 27 formal systems shows that ProGS improves over state-of-the-art agentic formal modeling approaches in syntactic validity, deductive verifiability, and behavioral correctness, demonstrating the benefit of organizing formal model construction around hierarchical proof sketches.