🤖 AI Summary
Existing benchmarks struggle to effectively evaluate the information retrieval and reasoning capabilities of large language models and agents in authentic scientific research settings. To address this gap, this work introduces a scientific exploration benchmark encompassing four task types across more than ten disciplines, comprising 103 expert-curated tasks. It establishes the first systematic framework for assessing progressive scientific information processing—from entity-level reasoning to cross-source knowledge synthesis—grounded in key technical pathways such as scientific database navigation, fuzzy literature retrieval, missing reference completion, and structured cross-source knowledge integration. Evaluation of over a dozen state-of-the-art models and agents reveals significant performance degradation on complex scientific tasks, with particularly low accuracy in structured knowledge synthesis.
📝 Abstract
Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level grounding and domain-level synthesis. We evaluate over ten state-of-the-art LLMs and autonomous agents on SciExplore, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks. These results highlight significant limitations of current models and agents in realistic scientific information-seeking scenarios.