From Tables to Quantified Statements: Evaluating LLM Inference Generation through Executable Verification

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过STAT-TO-TEXT任务评估大语言模型从统计表生成量化自然语言推理的能力,使用自动生成的Python代码验证推理准确性。
📝 Abstract
LLMs can generate fluent descriptions from tables, but their outputs may remain logically unsupported by the structured data. We introduce STAT-TO-TEXT, a controlled task in which LLMs generate quantified natural language inferences from statistical tables using quantified constructions such as all, some, no, and most. To evaluate these inferences, we use an LLM generated Python checker code which when executed verifies the corresponding truth conditions against the table. We compare four open-weight LLMs across model families and scales, evaluating faithfulness, logical accuracy, table coverage, and diversity. Our results show that model scale and family matter, with the largest model (GPT-OSS-120B) consistently producing the most faithful inferences without sacrificing greater table coverage and quantifier diversity, as opposed to smaller models. These findings are supported by human annotation, which shows that the automated checker closely aligns with human judgments.
Problem

Research questions and friction points this paper is trying to address.

LLMs
logical support
structured data
natural language inference
statistical tables
Innovation

Methods, ideas, or system contributions that make the work stand out.

quantified constructions
executable verification
automated checker
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mai Mohamed Eida
University of Illinois Urbana Champaign
G
Gunjan Anand
University of Illinois Urbana Champaign
Ayush Singh
Ayush Singh
Cigna, Northeastern University, Boston Children's Hospital, Harvard Medical School
Machine LearningDeep LearningComputer VisionNatural Language ProcessingBioInformatics
A
Aleksandre Maskharashvili
University of Illinois Urbana Champaign