Beyond Type-checking: Towards Holistic Evaluation of Formal Specification Generation
This study addresses the lack of deterministic verification of user intent in formal specification generation, which risks producing proofs grounded in erroneous specifications. To this end, it constructs a unified dataset and multidimensional evaluation framework encompassing formal validity, similarity, and behavioral adequacy. The work proposes a strategy distinguishing input acceptance from output constraints to clarify the evidential scope of metrics, and integrates an LLM agent workflow, the Lean theorem prover, and Generalized Tree Edit Distance (GTED) to enable automated evaluation. The findings reveal the limitations of single similarity metrics and the impact of metric coverage on ranking outcomes, demonstrating that perfect postcondition scores may obscure deficiencies in input contracts. Ultimately, this research provides a systematic benchmark for evaluating formal specifications.