BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing vision-and-language navigation platforms are tailored to domestic environments and fail to meet the precise navigation requirements of biomedical laboratories, particularly regarding instrument-proximal approach and safety clearance. This work proposes the first laboratory-oriented vision-and-language navigation simulation platform, introducing a novel three-zone target representation model—comprising the physical object, safety clearance zone, and operational zone—to uniformly support scene generation, task specification, navigation evaluation, and safety analysis. Combining procedural generation with manual design, the platform provides standardized interfaces and reinforcement learning compatibility for multimodal instruction-driven agent trajectory optimization. Experiments demonstrate that geometric exploration strategies achieve success rates of 74.4–87.5%, which further improve to 83.3–92.5% after multi-point sampling within operational zones, while substantially reducing the risk of unsafe proximity.
📝 Abstract
Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position. This representation is inadequate for laboratory instruments, which must be approached from their operating side while maintaining safe clearance from surrounding equipment. We introduce BioVLN, a simulation platform for developing and evaluating visual-language navigation agents in biomedical laboratories. BioVLN represents each instrument with three regions: its physical body, a surrounding clearance region, and an operation area in front of the usable side. This model is applied consistently to scene generation, target placement, navigation evaluation, and safety analysis, so success depends on reaching a position from which the instrument can be accessed. BioVLN supports procedural scene generation and manually designed environments, producing 47 scenes and 1667 episodes. Standardized navigation and reinforcement-learning interfaces enable trajectory collection and policy training. Experiments show that geometric exploration reaches 74.4--87.5% success, while sampling multiple valid positions in the operation area improves success to 83.3--92.5% and reduces unsafe proximity.
Problem

Research questions and friction points this paper is trying to address.

visual language navigation
biomedical laboratory
embodied navigation
instrument access
safe clearance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Language Navigation
Biomedical Laboratory Simulation
Three-Region Instrument Representation
Operation-Aware Navigation
Safety-Constrained Embodied AI
🔎 Similar Papers
No similar papers found.
Z
Zhe Liu
Key Laboratory of Smart Manufacturing in Energy Chemical Process, MoE, East China University of Science and Technology, Shanghai, China; Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China
Q
Quan Lu
Key Laboratory of Smart Manufacturing in Energy Chemical Process, MoE, East China University of Science and Technology, Shanghai, China; Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China
Z
Zhaohui Du
Key Laboratory of Smart Manufacturing in Energy Chemical Process, MoE, East China University of Science and Technology, Shanghai, China; Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China
Zhe Wang
Zhe Wang
Professor of Computer Science & Engineering, East China University of Science & Technology
Machine LearningPattern RecognitionMedical Data ProcessingImage AnalysisArtificial Intelligence
H
Huanbo Jin
Key Laboratory of Smart Manufacturing in Energy Chemical Process, MoE, East China University of Science and Technology, Shanghai, China; Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China
Jiaming Gu
Jiaming Gu
Institute of Automation, Chinese Academy of Sciences
Computer Vision
Qi Wang
Qi Wang
Shanghai Jiao Tong University << UCAS
Reinforcement LearningWorld ModelsComputer Vision
Ting Xiao
Ting Xiao
East China University of Science and Technology
Medical Image AnalysisFew-shot LearningReinforcement Learning
Minting Pan
Minting Pan
Shanghai Jiao Tong University
Machine Learning Reinforcement Learning
Dongzhan Zhou
Dongzhan Zhou
Researcher at Shanghai AI Lab
AI4Sciencecomputer visiondeep learning