DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?

๐Ÿ“… 2026-09-27
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the inadequacy of existing benchmarks in comprehensively evaluating AI agents across data validation, analytical verification, and evidence-driven hypothesis generation. To this end, we introduce DISCERN, a controlled evaluation framework constructed from real-world public datasets spanning eight life science domains. This benchmark uniquely integrates assessments of data integrity, reasoning rigor, and novelty through three-tiered testing and counterfactual scenarios, examining agentsโ€™ capacity to handle conflicting evidence and conduct scientific discovery under adversarial scrutiny. Comparative experiments across multiple models reveal that current AI agents perform poorly on higher-order hypothesis generation tasks, achieving merely 0.6% of the maximum score. These findings demonstrate that contemporary AI agents still lack reliable capabilities for autonomous analysis and scientific discovery.
๐Ÿ“ Abstract
Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical task completion or hypothesis generation separately rather than testing whether reliable evidence supports valid and novel claims. We introduce DISCERN (Data Integrity and Scientific Capability: Evidence, Reasoning, and Novelty), a controlled benchmark on real, publicly available datasets that evaluates three key levels of an automated research workflow. The first two levels test data integrity and analysis verification under confounds and tool traps, while the third tests hypothesis generation and revision under adversarial review, including counterfactual cases in which evidence consistent with real data and documented scientific phenomena conflicts with established expectations, motivating alternative explanations and testable hypotheses. Across 203 tasks, eight life-science tracks, and eight models, DISCERN shows that strong aggregate performance can mask level-specific weaknesses. Agents earn perfect scores in only 60.8% of Level 1, 34.2% of Level 2, and 0.6% of Level 3 evaluations, with penalties attributed to rejection of sound data, failure to carry recognized limitations into conclusions, and wide variation in hypothesis production. Cross-track rankings by token and code use are substantially more stable than rankings by evidence judgment, suggesting greater consistency in computational effort than in evidence-based reasoning. These profiles identify opportunities for supervised scientific assistance, but current agents do not yet demonstrate reliable autonomous analysis or discovery. Code and data: https://huggingface.co/datasets/discern-bench-anon/discern-benchmark
Problem

Research questions and friction points this paper is trying to address.

AI agents
automated research
scientific discovery
benchmark evaluation
hypothesis generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated Research Benchmark
AI Agents
Hypothesis Generation
Data Integrity
Adversarial Review
N
Nan Huang
University of California San Diego
M
Mario Tapia-Pacheco
University of California San Diego
K
Kun Zhou
University of California San Diego
Y
Yiming Huang
University of California San Diego
K
Kevin Josรฉ Barrientos Dรญaz
Independent Researcher
T
Tiffany Amariuta
University of California San Diego
Jingbo Shang
Jingbo Shang
Associate Professor, UC San Diego
Natural Language ProcessingData MiningDeep LearningInformation ExtractionWeak Supervision