BIABench: Evaluating AI agents on real-world bioimage analysis tasks

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of real-world, end-to-end evaluation benchmarks for AI agents in biological image analysis by constructing a 16-task benchmark encompassing multimodal challenges. Methodologically, it pioneers an evaluation framework integrating expert rule-based process scoring with standard outcome metrics to comprehensively assess both general-purpose and domain-specific AI agents across scenarios involving vision-language model judgment, code execution, and specialized software invocation. The findings reveal that while agents perform competently on conventional two-dimensional tasks, they encounter significant bottlenecks and exhibit pronounced instability in higher-dimensional tasks such as three-dimensional and temporal analyses. Ultimately, current methodologies remain insufficient to bridge this performance gap.
📝 Abstract
Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences are often too large to read as context, so an agent must choose and run an analysis through code, specialized software and rendered views. Published studies make this capability testable, because each pairs raw images with a peer-reviewed result. We introduce BIABench, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth. The tasks span eleven analysis subtasks and modalities from H&E histology to single-molecule localization microscopy. Each submission receives an outcome score, which compares the output files with the ground truth using field-standard metrics, and a process score, in which a vision-language model judges method choice and quality control against an expert-written rubric. We evaluated general-purpose and biology-specific agents across several language models, with repeated runs of every task. Routine two-dimensional tasks were solved well, but on some tasks that added a third dimension or a time axis no agent scored above 0.19. Neither biological specialization, stronger models nor detailed expert instructions closed this gap. The agents were also unreliable, with scores varying more between repeated runs of one agent than between different agents, and without ground truth a correct run could not be told from a wrong one by its process score or by the time spent. Released openly with its data and code, BIABench provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.
Problem

Research questions and friction points this paper is trying to address.

AI agents
bioimage analysis
benchmark
long-horizon tasks
reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bioimage analysis
AI agents benchmark
Vision-language model
Long-horizon tasks
Dual scoring mechanism
🔎 Similar Papers
No similar papers found.
Z
Zixuan Pan
Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, USA
D
Davide Panzeri
Leibniz-Institut für Analytische Wissenschaften – ISAS – e.V., Dortmund, Germany
L
Lukas Johanns
Leibniz-Institut für Analytische Wissenschaften – ISAS – e.V., Dortmund, Germany
M
Marilin Moor
Institute of Computer Science, University of Tartu, Tartu, Estonia
Yu Zhou
Yu Zhou
Leibniz-Institut für Analytische Wissenschaften - ISAS - e.V.
Deep Learning
Hedi Peterson
Hedi Peterson
Professor of Bioinformatics, University of Tartu; ELIXIR Estonia;
#unitartucsbioinformaticssoftware
Yiyu Shi
Yiyu Shi
Full Professor, University of Notre Dame
hardware/software co-designdeep learning accelerationon-device AIAI for healthcare
Jianxu Chen
Jianxu Chen
Group Leader, Leibniz-Institut für Analytische Wissenschaften – ISAS
Deep learning in biomedical image analysis and computer Vision