StatEval: A Comprehensive Benchmark for Large Language Models in Statistics

πŸ“… 2025-10-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF

career value

180K/year
πŸ€– AI Summary
The statistics domain lacks a systematic, large language model (LLM)-oriented evaluation benchmark. Method: We introduce StatEvalβ€”the first comprehensive benchmark for statistical reasoning, covering undergraduate and graduate curricula as well as frontier research topics, comprising 13,817 foundational questions and 2,374 research-level formal proof tasks. We design a scalable multi-agent automated pipeline integrated with human verification to ensure question quality, and develop a fine-grained evaluation framework tailored to statistical computation and formal proof. Contribution/Results: Experiments reveal severe limitations of state-of-the-art LLMs on research-level statistical tasks (e.g., GPT-5-mini achieves <57% accuracy; open-source models perform worse), highlighting the unique challenges of statistical reasoning. StatEval fills a critical gap in the field and provides a rigorous, reliable benchmark for diagnostic model assessment and algorithmic advancement.

Technology Category

Application Category

πŸ“ Abstract
Large language models (LLMs) have demonstrated remarkable advances in mathematical and logical reasoning, yet statistics, as a distinct and integrative discipline, remains underexplored in benchmarking efforts. To address this gap, we introduce extbf{StatEval}, the first comprehensive benchmark dedicated to statistics, spanning both breadth and depth across difficulty levels. StatEval consists of 13,817 foundational problems covering undergraduate and graduate curricula, together with 2374 research-level proof tasks extracted from leading journals. To construct the benchmark, we design a scalable multi-agent pipeline with human-in-the-loop validation that automates large-scale problem extraction, rewriting, and quality control, while ensuring academic rigor. We further propose a robust evaluation framework tailored to both computational and proof-based tasks, enabling fine-grained assessment of reasoning ability. Experimental results reveal that while closed-source models such as GPT5-mini achieve below 57% on research-level problems, with open-source models performing significantly lower. These findings highlight the unique challenges of statistical reasoning and the limitations of current LLMs. We expect StatEval to serve as a rigorous benchmark for advancing statistical intelligence in large language models. All data and code are available on our web platform: https://stateval.github.io/.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLMs' statistical reasoning across foundational and research-level problems
Addressing the lack of comprehensive benchmarks for statistical intelligence in AI
Assessing performance gaps in LLMs for computational and proof-based statistical tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated multi-agent pipeline for problem extraction
Human-in-the-loop validation ensuring academic rigor
Tailored evaluation framework for statistical reasoning
πŸ”Ž Similar Papers
No similar papers found.
Y
Yuchen Lu
Shanghai University of Finance and Economics
Run Yang
Run Yang
Harbin Institute of Technology
edge computingartificial intelligencenetwork security
Y
Yichen Zhang
Shanghai University of Finance and Economics
S
Shuguang Yu
Shanghai University of Finance and Economics
R
Runpeng Dai
University of North Carolina at Chapel Hill
Z
Ziwei Wang
Shanghai University of Finance and Economics
J
Jiayi Xiang
Shanghai University of Finance and Economics
W
Wenxin E
Shanghai University of Finance and Economics
S
Siran Gao
Shanghai University of Finance and Economics
X
Xinyao Ruan
Shanghai University of Finance and Economics
Y
Yirui Huang
Shanghai University of Finance and Economics
C
Chenjing Xi
Shanghai University of Finance and Economics
H
Haibo Hu
Shanghai University of Finance and Economics
Y
Yueming Fu
Shanghai University of Finance and Economics
Q
Qinglan Yu
Shanghai University of Finance and Economics
X
Xiaobing Wei
Shanghai University of Finance and Economics
J
Jiani Gu
Shanghai University of Finance and Economics
R
Rui Sun
Shanghai University of Finance and Economics
J
Jiaxuan Jia
Shanghai University of Finance and Economics
F
Fan Zhou
Shanghai University of Finance and Economics