Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery

πŸ“… 2026-02-10
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Current static evaluation benchmarks are vulnerable to training data contamination and thus struggle to assess AI systems’ capacity for discovering novel knowledge. To address this limitation, this work proposes DBench-Bioβ€”the first dynamic, fully automated, and monthly updated benchmark for biomedical knowledge discovery, spanning twelve subfields. The framework continuously ingests recent authoritative paper abstracts and leverages large language models to generate hypothetical question-answer pairs, which undergo multidimensional quality filtering based on relevance, clarity, and centrality to ensure novelty and freedom from contamination. Evaluation of state-of-the-art models using DBench-Bio reveals significant deficiencies in their ability to reason over newly emerging knowledge, while simultaneously offering the research community its first sustainably evolving benchmark resource for ongoing assessment and development.
πŸ“ Abstract
Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery. However, rigorously evaluating an AI's capacity for knowledge discovery remains a critical challenge. Existing benchmarks predominantly rely on static datasets, leading to inevitable data contamination where models have likely seen the evaluation knowledge during training. Furthermore, the rapid release cycles of modern LLMs render static benchmarks quickly outdated, failing to assess the ability to discover truly new knowledge. To address these limitations, we propose DBench-Bio, a dynamic and fully automated benchmark designed to evaluate AI's biological knowledge discovery ability. DBench-Bio employs a three-stage pipeline: (1) data acquisition of rigorous, authoritative paper abstracts; (2) QA extraction utilizing LLMs to synthesize scientific hypothesis questions and corresponding discovery answers; and (3) QA filter to ensure quality based on relevance, clarity, and centrality. We instantiate this pipeline to construct a monthly-updated benchmark covering 12 biomedical sub-domains. Extensive evaluations of SOTA models reveal current limitations in discovering new knowledge. Our work provides the first dynamic, automatic framework for assessing the new knowledge discovery capabilities of AI systems, establishing a living, evolving resource for AI research community to catalyze the development of knowledge discovery.
Problem

Research questions and friction points this paper is trying to address.

knowledge discovery
large language models
dynamic benchmark
biological knowledge
evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

dynamic benchmark
knowledge discovery
large language models
automated evaluation
biomedical QA
πŸ”Ž Similar Papers
No similar papers found.