🤖 AI Summary
Low trustworthiness, difficulty in error diagnosis, and weak human-AI collaboration arise from the “black-box” nature of machine learning in physics. Method: This study systematically constructs the first multidimensional interpretability classification framework tailored to physical sciences, integrating philosophical reflection with technical practice. It unifies diverse interpretability approaches—including surrogate modeling, attention mechanisms, symbolic regression, causal inference, and physics-informed constraint embedding—across condensed matter, high-energy, astrophysical, and statistical physics. Contribution/Results: We identify a fundamental interpretability–performance trade-off law and establish verifiable, reproducible evaluation metrics and application guidelines. The framework elevates interpretable AI from a mere analytical tool to a core paradigm of scientific intelligence, enabling human-understandable, automated scientific discovery.
📝 Abstract
Machine learning is increasingly transforming various scientific fields, enabled by advancements in computational power and access to large data sets from experiments and simulations. As artificial intelligence (AI) continues to grow in capability, these algorithms will enable many scientific discoveries beyond human capabilities. Since the primary goal of science is to understand the world around us, fully leveraging machine learning in scientific discovery requires models that are interpretable -- allowing experts to comprehend the concepts underlying machine-learned predictions. Successful interpretations increase trust in black-box methods, help reduce errors, allow for the improvement of the underlying models, enhance human-AI collaboration, and ultimately enable fully automated scientific discoveries that remain understandable to human scientists. This review examines the role of interpretability in machine learning applied to physics. We categorize different aspects of interpretability, discuss machine learning models in terms of both interpretability and performance, and explore the philosophical implications of interpretability in scientific inquiry. Additionally, we highlight recent advances in interpretable machine learning across many subfields of physics. By bridging boundaries between disciplines -- each with its own unique insights and challenges -- we aim to establish interpretable machine learning as a core research focus in science.