SleuthBench: Benchmarking Statistical LLM Evaluation Using Tabular Hidden Signals

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost of ground-truth acquisition and model memorization interference in evaluating the statistical discovery capabilities of large language models. To mitigate these issues, this work introduces a novel automated benchmarking framework based on hidden signal injection, which embeds controllable patterns into publicly available tabular data and automatically computes reference answers to circumvent memorization effects. Furthermore, it proposes an "empirical layer" composed of precomputed statistical artifacts, integrated with code-based tooling and interaction effect fitting techniques to enhance feature contribution analysis. Experimental results demonstrate that the proposed approach achieves 83.8% accuracy in data quality detection. Notably, incorporating the empirical layer substantially improves feature contribution accuracy from 41.9% to 68.0%, validating the effectiveness of the framework for rigorous LLM evaluation.
📝 Abstract
Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introduce SLEUTHBENCH, a benchmark that addresses both problems by injecting controlled data-quality problems and feature effects into public tabular datasets: the injected pattern determines the answer, so reference answers are computed automatically and memorized knowledge of the original table is insufficient, while the table keeps its background structure. The injected patterns are modeled on phenomena reported in real data analyses. The benchmark defines 17 question templates in two families: data-quality questions and feature-contribution questions. We evaluate six state-of-the-art LLMs that analyze the data using a Python coding tool, on data-science and business phrasings of 70 validated dataset-template combinations, yielding 1680 graded responses in total. The models detect data-quality problems reliably (83.8% accuracy) but recover feature contributions poorly (41.9%). Finding how features shape the target requires searching over both candidate variables and analytical procedures. To address this issue, we propose the Empirical Layer, a set of precomputed statistical artifacts comprising summaries, fitted feature and interaction effects, and dataset descriptions, which exposes candidate patterns for direct inspection. Access to these artifacts raises feature-contribution accuracy from 41.9% to 68.0%.
Problem

Research questions and friction points this paper is trying to address.

LLM evaluation
tabular data
statistical discovery
benchmark
data quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

SleuthBench
Tabular Hidden Signals
Empirical Layer
Statistical LLM Evaluation
Feature Contribution
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jingyun Jia
University of Wisconsin–Madison, Intelligible
Antoine Remond-Tiedrez
Antoine Remond-Tiedrez
Intelligible
A
Aaron Alvarez
University of Cincinnati
J
Joshua Shunk
Intelligible
Rich Caruana
Rich Caruana
Microsoft Research
machine learning
Ben Lengerich
Ben Lengerich
University of Wisconsin-Madison
MLAIMedical InformaticsComputational Genomics