RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of standardized difficulty stratification and automated verification in existing benchmarks by constructing a dynamic benchmark comprising 100 NeurIPS papers, introducing a novel three-tier difficulty classification system based on released resources. Methodologically, large language model-driven agents are employed to execute code debugging, retraining, and rewriting tasks, with an independent language model introduced for automated scoring. Results demonstrate that the best-performing agent achieves reproduction rates of merely 15% to 41% across difficulty levels, with most failures attributable to unverified numerical results and an average consumption of only 29% of the allocated computational budget. This work establishes an annually updatable evaluation paradigm and reveals significant limitations of current AI agents in scientific reproduction.
📝 Abstract
Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the authors released decides the difficulty tier. Run-tier releases include code, data, and weights; Retrain-tier releases lack weights, so the agent trains the model; Reimplement-tier releases lack code, so the agent writes it. A separate language model grades runs from logs and outputs rather than agents' reports. We run four agents once per paper; the best agent in each tier reproduces only 41% of Run-tier papers, 27% at Retrain, and 15% at Reimplement, where every agent does worst. Failed attempts use on average 29% of their budget, so most stop with budget left. The most common agent error is writing the method without checking any part against the paper's numbers, in 63 of 400 runs.
Problem

Research questions and friction points this paper is trying to address.

AI agents
reproducibility
machine learning papers
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

benchmark
AI agents
reproducibility
difficulty tiers
LLM evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mithil Salunkhe
University of Illinois Urbana-Champaign
H
Haochen Ding
University of Illinois Urbana-Champaign
S
Samridhi Verma
University of Illinois Urbana-Champaign
Volodymyr Kindratenko
Volodymyr Kindratenko
University of Illinois at Urbana-Champaign
HPCAI