EnVisionVR: A Scene Interpretation Tool for Visual Accessibility in Virtual Reality

📅 2025-02-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenges faced by blind and low-vision (BLV) users in acquiring visual information, comprehending 3D virtual environments, and interacting with virtual objects in VR, this paper proposes the first end-cloud collaborative Vision-Language Model (VLM)-based system for VR accessibility. The system enables plug-and-play, real-time visual-semantic interpretation without modifying existing VR applications. It integrates speech interaction, spatial audio feedback, and lightweight 3D scene semantic parsing, dynamically distributing computational load between the VR client and cloud. Crucially, it pioneers the integration of VLMs into VR accessibility frameworks, balancing real-time performance, natural language interaction, and multimodal accessibility. In an evaluation involving 12 BLV participants, object localization accuracy improved by 67%, and 92% reported significant enhancements in scene understanding and interactive autonomy.

Technology Category

Computer Vision: Language and VisionHumans and AI: AI for AccessibilityMachine Learning: Large Multimodal Models (LMMs)

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSystems and Infrastructure for Web, Mobile and WoT: Virtualization and resource management in Web systems and infrastructures
📝 Abstract
Effective visual accessibility in Virtual Reality (VR) is crucial for Blind and Low Vision (BLV) users. However, designing visual accessibility systems is challenging due to the complexity of 3D VR environments and the need for techniques that can be easily retrofitted into existing applications. While prior work has studied how to enhance or translate visual information, the advancement of Vision Language Models (VLMs) provides an exciting opportunity to advance the scene interpretation capability of current systems. This paper presents EnVisionVR, an accessibility tool for VR scene interpretation. Through a formative study of usability barriers, we confirmed the lack of visual accessibility features as a key barrier for BLV users of VR content and applications. In response, we designed and developed EnVisionVR, a novel visual accessibility system leveraging a VLM, voice input and multimodal feedback for scene interpretation and virtual object interaction in VR. An evaluation with 12 BLV users demonstrated that EnVisionVR significantly improved their ability to locate virtual objects, effectively supporting scene understanding and object interaction.
Problem

Research questions and friction points this paper is trying to address.

Enhances visual accessibility in VR
Leverages VLM for scene interpretation
Supports BLV users in VR interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Leverages Vision Language Models
Uses voice input
Multimodal feedback system