🤖 AI Summary
This study investigates whether frontier AI models can reliably curate training data and predict fine-tuning performance. To this end, we introduce a novel benchmark requiring AI agents to select and rank optimal data subsets for large language model fine-tuning from candidate pools through code analysis and forward reasoning. This work provides the first quantitative evaluation of AI “data awareness,” revealing inconsistencies in interpreting data signals via execution trajectory analysis. Experimental results demonstrate that agent-driven data selection yields limited performance gains, with unstable ranking accuracy and insufficient cross-task generalization. These findings offer empirical evidence for understanding the data selection bottlenecks inherent in current AI agents.
📝 Abstract
As claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data for training? We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning. We ask AI agents to select and rank candidate training subsets that can be used to fine-tune a small LLM model. Agents are allowed to inspect the data, write and execute analysis code, and run model forward passes, but can not train the model or access the actual evaluation tasks. We then fine-tune the base model on each selected subset and evaluate its post-training performance under a standardized protocol. We instantiate the benchmark in terminal problem solving and tool use, selecting trajectories from OpenThoughts-Agent and EnvScaler and evaluating on TBLite and BFCL, respectively. We then evaluate the agents along two complementary dimensions: the post-training performance of the top-ranked subset, reflecting the ability to identify high-value training data, and ranking accuracy, reflecting the ability to predict the relative performance of the selected subsets. In our experiments, selection gains over random selection are limited; agents do not reliably rank their selected groups, and ranking ability does not hold consistently across tasks: Astra identifies the best group in all three tool-use runs but in only one of three terminal runs. Analysis of execution traces on both tasks shows that agents often use similar data signals while interpreting their training value differently.