OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the absence of evaluation benchmarks for assessing AI capabilities in solving fundamental open problems in science by constructing a benchmark comprising 82 unsolved challenges in mathematics and physics. Methodologically, it introduces a novel evaluation framework grounded in authentic scientific literature that operates without predefined ground-truth answers, alongside a multi-evaluator model ensemble and a problem-context modeling mechanism to objectively quantify solution progress. Experimental results demonstrate that GPT-6-Astra achieves the highest resolution rate of 14.0%, significantly outperforming existing open-source and lightweight models. By bridging the gap in evaluating AI-driven frontier scientific exploration, this work establishes a reliable paradigm for assessing the reasoning limits of large language models on open-ended problems.
πŸ“ Abstract
The next frontier for artificial general intelligence is tackling unresolved scientific problems, calling for benchmarks that assess progress beyond established knowledge. We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature. Each problem supplies the research context, assumptions, and prior progress needed to investigate the question. We select problems whose proposed solutions admit comparatively clear checks of their decisive mathematical or computational claims. Four evaluator models independently assess the correctness, completeness, and degree of progress of each submission without reference solutions. Across seven evaluated configurations, GPT-6-Astra achieves the highest mean judged solve rate of 14.0%, compared with 5.5-6.7% for the evaluated full-size open models and 2.4-3.7% for Flash models. Case comparisons connect stronger outcomes to changes in problem representation, general arguments that extend beyond finite evidence, and proofs of the steps needed to complete a solution. By grounding evaluation in questions arising from the research literature, OpenProblemBench provides a setting for investigating the capabilities and limitations of AI as a contributor to foundational theoretical science.
Problem

Research questions and friction points this paper is trying to address.

Open problems
Artificial general intelligence
Benchmarking
Theoretical sciences
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

OpenProblemBench
benchmarking
open problems
theoretical sciences
LLM evaluation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Zhiyi Li
Zhiyi Li
University of Science and Technology of China
Statistical mechanics, Quantum many body physics
S
Sihan Hu
Department of Modern Physics, University of Science and Technology of China, Hefei, Anhui 230026, China
T
Tianning Xiao
Hefei National Laboratory, University of Science and Technology of China, Hefei 230088, China
Xiansheng Cai
Xiansheng Cai
Institute of Theoretical Physics, CAS
Monte Carloeffective field theorysuperconductivitymachine learning
X
Xiaojun Tan
Institute for Advanced Algorithms Research, Shanghai 200120, China
Youjin Deng
Youjin Deng
University of Science and Technology of China
Computational Statistical Physics and Condensed-Matter Physics
K
Kun Chen
CAS Key Laboratory of Theoretical Physics, Institute of Theoretical Physics, Chinese Academy of Sciences, Beijing 100190, China