Quizzing the Translation: A Prover-Grounded Evaluation Metric for NL$\rightarrow$FOL

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of traditional metrics in detecting severe logical errors during the evaluation of natural language to first-order logic translation. To this end, it proposes the SIV metric, which pioneers the integration of automated theorem proving into the evaluation framework. By generating positive and negative contrastive probes to verify the logical consistency of candidate translations, SIV precisely distinguishes omission from over-assertion errors, enabling interpretable error tracing. Experimental results demonstrate that SIV achieves an error correlation of 80% and a correct ranking rate exceeding 99%. Furthermore, it attains a significantly superior AUC in real-world large language model translation evaluations, effectively stratifying critical logical errors.
📝 Abstract
A standard pipeline for symbolic reasoning over natural-language problems translates them into first-order logic and invokes a theorem prover. The translation step is the bottleneck: swap"every"for"some"and every inference that follows is corrupted. Yet today's metrics often score more broken translations higher than less broken ones, because BLEU, BERTScore, and Smatch++ reward surface overlap that the worst errors happen to preserve. We introduce SIV, which derives two kinds of probes from the target formula and uses a theorem prover to verify the candidate translation against each. Positive probes are statements the candidate must entail, which detect translations that drop content; contrastive probes are statements the candidate must not entail, which detect translations that assert more than the original. On a controlled pool of perturbed FOLIO translations, the severity of the error accounts for 80% of SIV's score variance, compared with at most 17% for any prior metric. Across six error classes on a disjoint pool, SIV scores the reference above the perturbed candidate in over 99% of pairs. Because each probe is labeled with what it tests, the failure pattern also supplies a labeled error trace, recovering the perturbation class at macro-F1 0.638, nearly double the score-only baseline. On 434 expert-audited real LLM translations, SIV attains the top AUC, uniquely detects and grades expert-labeled major errors, and abstains, rather than mis-scoring, on out-of-vocabulary translations.
Problem

Research questions and friction points this paper is trying to address.

Natural Language to First-Order Logic
Translation Evaluation Metric
Symbolic Reasoning
Semantic Error Detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluation Metric
First-Order Logic
Theorem Prover
Symbolic Reasoning
Contrastive Probes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.