π€ AI Summary
This study addresses the limitation of existing benchmarks in evaluating AI agentsβ capacity to construct scientific models when no ground-truth solutions are available. To this end, it introduces the ENSO climate modeling benchmark, which requires large language model agents to autonomously develop low-order stochastic models from real-world observational data within a time constraint, validated through a concealed multidimensional scoring mechanism assessing statistical reproducibility and predictive skill. This work achieves the first evaluation of scientific modeling without predefined answers. Experiments demonstrate that six out of twelve agents outperform previously published models and independently align with unmentioned contested theories of ENSO, confirming that AI can generate competitive scientific hypotheses and construct effective climate models.
π Abstract
Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations. Within a six-hour budget, agents process the observations, write their own diagnostics, which are then frozen, and develop a model using only these diagnostics as feedback. Hidden graders then test whether the model reproduces ENSO's statistics, recovers unobserved variables, and forecasts held-out years, and score a published model in the same way. Across twelve agent systems, six produce models that score higher than the published model, mainly through better reconstruction and forecasting. The simplified forms of the stronger models are each compatible with one of the two competing explanations of ENSO's warm-cold asymmetry, an open debate that the task never mentions. Controlled runs of the top system under varied information suggest that its scores do not come from recalling the dated observational record and that the information it receives shapes how it builds its model. SciExam for ENSO can thus evaluate agent research where no answer is known, and the results suggest that agents can already build competitive models whose structures bear on questions that scientists still debate.