OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing AI virtual cell benchmarks focus exclusively on simulation prediction, lacking evaluation of models' capacity to interpret evidence and generate biological hypotheses. To address this limitation, this work introduces OmniVCBench, a chart-centric and traceable benchmark comprising 6,077 question-answer pairs spanning three reasoning tasks: prediction, explanation, and discovery. Furthermore, it proposes AIVC-Judge, a multimodal evaluation framework that integrates Bloom’s taxonomy with hard negative mining strategies to enable cost-effective assessment of open-ended responses. Experimental results demonstrate a significant positive correlation between multiple-choice accuracy and judge scores, establishing a complementary and reliable paradigm for evaluating the scientific reasoning capabilities of AI virtual cells.
📝 Abstract
Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question--answer pairs derived from figures and experimental contexts in the scientific literature. Guided by Bloom's taxonomy, we instantiate interpretation-layer counterparts of the AIVC Predict--Explain--Discover agenda through three scientific reasoning tasks. We further introduce AIVC-Judge, a task-conditioned MLLM-as-a-judge framework with category-specific, reference-aware rubrics for evaluating open-ended responses. A complementary Model-Derived Hard-Negative Mining (MDHNM) strategy converts plausible errors observed during model inference into MCQ distractors for lower-cost evaluation. Within the evaluated heterogeneous model pool, MCQ accuracy correlates positively with AIVC-Judge scores, providing a complementary view of performance alongside open-response evaluation. Code and data demo are available at https://anonymous.4open.science/r/OmniVCBench.
Problem

Research questions and friction points this paper is trying to address.

Artificial Intelligence Virtual Cells
Multimodal Reasoning
Benchmark
Scientific Evidence Interpretation
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Reasoning Benchmark
MLLM-as-a-Judge
Hard-Negative Mining
AI Virtual Cells
Scientific Reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.