BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing benchmarks that prioritize factual recall over reasoning and struggle to evaluate frontier biological knowledge and multimodal capabilities. To this end, we construct a global, multi-institutional, doctoral-level evaluation benchmark for bioengineering spanning eleven subfields. This benchmark introduces novel interdisciplinary tasks—encompassing multiple-choice, literature synthesis, and multimodal formats—that specifically target experimental reasoning and multimodal interpretation, supported by an expert blind-review consensus mechanism for quality control. Leveraging a large language model evaluation framework with comparative analyses of cloud-based and local deployments, the best-performing model achieves 90% accuracy and a similarity score of 0.72. Our findings reveal significant performance disparities across subfields and establish a scalable, standardized evaluation protocol for assessing advanced AI capabilities in bioengineering.
📝 Abstract
Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and multimodal tasks. We assembled BioEVAL (BioEngineering Validation of AI and LLMs), a global, multi-institutional initiative designed to assess experimental reasoning capability across bioengineering (BE) subfields. BioEVAL spans 11 major BE subfields plus a set of uncategorized items, bringing together 22 research groups to create a PhD-level benchmark comprising 608 evaluation items: 1) 380 multiple-choice questions (MCQs, 359 retained after audit), 2) 218 literature synthesis tasks, and 3) 10 multimodal problems with experimental image interpretation. Benchmark items underwent authoring-group expert review and centralized quality control before evaluation. Following evaluation, a blinded cross-group consensus audit of the highest- and lowest-accuracy MCQ items flagged 21 questions for revision or removal; these were withheld, and all reported MCQ results are computed on the 359 retained items. We evaluated diverse cloud-scale foundation/multimodal models (e.g., ChatGPT, Gemini, and Grok) and locally deployable models suitable for inference on consumer-grade GPUs. Models achieved the highest accuracy of up to 90% on MCQs, similarity score of 0.72 on literature synthesis, and accuracy of 80% on a small sample of multimodal reasoning questions, with substantial performance variation across subfields. Leaderboard rankings characterize current capabilities, limitations, and development priorities across the evaluated BE task categories. BioEVAL is maintained as an extensible benchmark with standardized protocols for continuing expert item contribution and model evaluation.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Bioengineering
Benchmarking
Multimodal reasoning
Experimental reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmark
Large Language Models
Multimodal Reasoning
Bioengineering
Experimental Reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shun Ye
Department of Bioengineering, University of California, Los Angeles, CA 90095, USA
V
Vinny Chandran Suja
Harvard John A. Paulson School of Engineering and Applied Sciences, Harvard University, Allston, MA 02134, USA
C
Chenlong Li
Department of Bioengineering, University of California, Los Angeles, CA 90095, USA
C
Chongming Jiang
Terasaki Institute for Biomedical Innovation, Woodland Hills, CA 91367, USA
R
Reza Zamani
Terasaki Institute for Biomedical Innovation, Woodland Hills, CA 91367, USA
X
Xiang Li
Department of Bioengineering, University of California, Los Angeles, CA 90095, USA
C
Christopher Bain
The Wallace H. Coulter Department of Biomedical Engineering, Georgia Institute of Technology, Atlanta, GA 30332, USA
Y
Yuqi Zhou
Department of Chemistry, The University of Tokyo, Tokyo 113-0033, Japan
W
Walker Peterson
Department of Chemistry, The University of Tokyo, Tokyo 113-0033, Japan
H
Huidong Wang
Department of Chemistry, The University of Tokyo, Tokyo 113-0033, Japan
C
Chenglang Hu
Department of Chemistry, The University of Tokyo, Tokyo 113-0033, Japan
Jongchan Park
Jongchan Park
Lunit Inc.
X
Xiao Cheng
Department of Biomedical Engineering, Columbia University, New York, NY 10027, USA
B
Benjamin Swedlund
Eli and Edythe Broad CIRM Center for Regenerative Medicine and Stem Cell Research, Keck School of Medicine, University of Southern California, Los Angeles, CA 90033, USA
S
Sandra Murillo
Eli and Edythe Broad CIRM Center for Regenerative Medicine and Stem Cell Research, Keck School of Medicine, University of Southern California, Los Angeles, CA 90033, USA
A
Anjali Sivanandan
Department of Bioengineering, University of California, Los Angeles, CA 90095, USA
Shiyu Sun
Shiyu Sun
George Mason University
Data MiningSoftware SecurityProgram Analysis
L
Liang Lanfeng
Mechanobiology Institute, National University of Singapore, Singapore 117411, Singapore
M
Mohammad Tariqul Islam
Nano-Cybernetic Biotrek, Media Lab, Massachusetts Institute of Technology, Cambridge, MA 02139, USA
B
Baju C. Joy
Nano-Cybernetic Biotrek, Media Lab, Massachusetts Institute of Technology, Cambridge, MA 02139, USA
I
Ishaq N. Khan
Nano-Cybernetic Biotrek, Media Lab, Massachusetts Institute of Technology, Cambridge, MA 02139, USA
S
Sreedhar S. Kumar
Department of Biosystems Science and Engineering, ETH Zurich, CH-4056 Basel, Switzerland
G
Gabriel Mercado-Vásquez
Pritzker School of Molecular Engineering, University of Chicago, Chicago, IL 60637, USA
J
James V. Vizzard
Pritzker School of Molecular Engineering, University of Chicago, Chicago, IL 60637, USA
Jonathan M. Matthews
Jonathan M. Matthews
Department of Medicine, Biological Sciences Division, University of Chicago, Chicago, IL 60637, USA