Reasoning With a Star: A Heliophysics Dataset and Benchmark for Agentic Scientific Reasoning

📅 2025-11-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) lack physics-driven scientific reasoning capabilities in solar physics. Method: We introduce SunPhysBench—the first structured scientific reasoning benchmark for heliophysics—built from NASA/UCAR summer school problem sets, featuring unit-aware, hypothesis-explicit, and format-standardized question-answer pairs. We decompose reasoning tasks using systems engineering principles and design a programmatic evaluator supporting unit consistency checking, symbolic equivalence matching, and pattern validation. We compare single-prompt baselines against four multi-agent collaborative workflows. Contribution/Results: Experimental results demonstrate that multi-agent decomposition-based reasoning significantly outperforms baselines on deductive tasks, validating that structured collaboration enhances LLMs’ physical reasoning fidelity. SunPhysBench provides a rigorous, domain-specific evaluation framework to advance physics-informed AI for solar and space science.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Knowledge Representation and Reasoning: Computational Complexity of ReasoningMultiagent Systems: Teamwork

Application Category

Search and Retrieval-Augmented AI: Agentic searchEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Scientific reasoning through Large Language Models in heliophysics involves more than just recalling facts: it requires incorporating physical assumptions, maintaining consistent units, and providing clear scientific formats through coordinated approaches. To address these challenges, we present Reasoning With a Star, a newly contributed heliophysics dataset applicable to reasoning; we also provide an initial benchmarking approach. Our data are constructed from National Aeronautics and Space Administration & University Corporation for Atmospheric Research Living With a Star summer school problem sets and compiled into a readily consumable question-and-answer structure with question contexts, reasoning steps, expected answer type, ground-truth targets, format hints, and metadata. A programmatic grader checks the predictions using unit-aware numerical tolerance, symbolic equivalence, and schema validation. We benchmark a single-shot baseline and four multi-agent patterns, finding that decomposing workflows through systems engineering principles outperforms direct prompting on problems requiring deductive reasoning rather than pure inductive recall.
Problem

Research questions and friction points this paper is trying to address.

Advancing scientific reasoning in heliophysics beyond factual recall
Developing benchmark for agentic reasoning with physical assumptions and units
Evaluating multi-agent patterns for deductive heliophysics problem-solving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent reasoning patterns for scientific deduction
Programmatic grader with unit-aware validation
Structured dataset with physics problem decomposition
K
Kevin Lee
Frontier Development Lab, Frisco, TX, USA; Department of Mechanical and Aerospace Engineering, UCLA, Los Angeles, CA, USA
R
Russell Spiewak
Frontier Development Lab, Frisco, TX, USA; Trillium Technologies Inc., Frisco, TX, USA
James Walsh
James Walsh
Enterprise Fellow/Senior Lecturer, University of South Australia
ARVRTangible user interfacesHCI