AutoDataBench: A Data-centric Testbed for Accelerating Auto Research

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of existing benchmarks in attributing performance differences among agents and their lack of independent evaluation for data intelligence. To this end, this work proposes a conceptual framework for "data intelligence" and constructs a testbed that isolates non-data factors. It pioneers the integration of predictive and feedback mechanisms to validate the iterative optimization and causal reasoning capabilities of large language models (LLMs) in data diagnosis, organization, and construction. Technically, the approach synthesizes LLM agents, tool invocation, retrieval augmentation, knowledge injection, and iterative experimentation. Empirical results demonstrate that this framework effectively quantifies data intelligence. Furthermore, reusing experimental trajectories for intermediate training significantly enhances downstream code generation performance.
πŸ“ Abstract
Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at https://github.com/AutoDataBench/AutoDataBench.
Problem

Research questions and friction points this paper is trying to address.

Auto-research
Data Intelligence
Benchmark
Large Language Models
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Intelligence
AutoDataBench
Iterative Experimentation
Data-effect Reasoning
Mid-training
πŸ”Ž Similar Papers
No similar papers found.