OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences
This study addresses the absence of evaluation benchmarks for assessing AI capabilities in solving fundamental open problems in science by constructing a benchmark comprising 82 unsolved challenges in mathematics and physics. Methodologically, it introduces a novel evaluation framework grounded in authentic scientific literature that operates without predefined ground-truth answers, alongside a multi-evaluator model ensemble and a problem-context modeling mechanism to objectively quantify solution progress. Experimental results demonstrate that GPT-6-Astra achieves the highest resolution rate of 14.0%, significantly outperforming existing open-source and lightweight models. By bridging the gap in evaluating AI-driven frontier scientific exploration, this work establishes a reliable paradigm for assessing the reasoning limits of large language models on open-ended problems.