🤖 AI Summary
Large language models (LLMs) lack physics-driven scientific reasoning capabilities in solar physics. Method: We introduce SunPhysBench—the first structured scientific reasoning benchmark for heliophysics—built from NASA/UCAR summer school problem sets, featuring unit-aware, hypothesis-explicit, and format-standardized question-answer pairs. We decompose reasoning tasks using systems engineering principles and design a programmatic evaluator supporting unit consistency checking, symbolic equivalence matching, and pattern validation. We compare single-prompt baselines against four multi-agent collaborative workflows. Contribution/Results: Experimental results demonstrate that multi-agent decomposition-based reasoning significantly outperforms baselines on deductive tasks, validating that structured collaboration enhances LLMs’ physical reasoning fidelity. SunPhysBench provides a rigorous, domain-specific evaluation framework to advance physics-informed AI for solar and space science.
📝 Abstract
Scientific reasoning through Large Language Models in heliophysics involves more than just recalling facts: it requires incorporating physical assumptions, maintaining consistent units, and providing clear scientific formats through coordinated approaches. To address these challenges, we present Reasoning With a Star, a newly contributed heliophysics dataset applicable to reasoning; we also provide an initial benchmarking approach. Our data are constructed from National Aeronautics and Space Administration & University Corporation for Atmospheric Research Living With a Star summer school problem sets and compiled into a readily consumable question-and-answer structure with question contexts, reasoning steps, expected answer type, ground-truth targets, format hints, and metadata. A programmatic grader checks the predictions using unit-aware numerical tolerance, symbolic equivalence, and schema validation. We benchmark a single-shot baseline and four multi-agent patterns, finding that decomposing workflows through systems engineering principles outperforms direct prompting on problems requiring deductive reasoning rather than pure inductive recall.