Score
Designing input mappings and multimodal (haptic, audio, AR) feedback strategies to convey scene information, enable embodied interactions (e.g., animal embodiment in VR), and guide visually impaired users to detected objects.
Existing surveys on visual multimodal interfaces (VMIs) are predominantly task- or scenario-oriented, lacking a unified design paradigm. Method: This paper proposes a novel, vision-anchored, data-modality-driven classification and design framework, established through systematic literature review and cross-dimensional modeling. It introduces a four-dimensional taxonomy encompassing input modalities, fusion mechanisms, interaction objectives, and deployment environments, structured hierarchically as “holistic–detail–holistic.” Contribution/Results: The framework transcends conventional survey limitations by positioning the visual modality at the core of context-aware system design for the first time, integrating theories from human-computer interaction, multimodal learning, and context modeling. It delivers a reusable design methodology and principled guidelines for high-fidelity user intent understanding and seamless physical-digital interaction.
This study addresses the challenges of reduced precision and user confidence in optical see-through augmented reality (AR) during handheld tool guidance, which are often caused by visual occlusion, illumination variations, and ambiguous interface cues. To overcome these limitations, the authors propose a multimodal guidance system that integrates AR visual feedback with wrist-worn haptic feedback, introducing for the first time directional and state-based vibrotactile cues tailored for surgical-grade tool manipulation. The system incorporates a reference mapping mechanism informed by surgeon preferences, a custom wrist-mounted haptic device, and a user-centered vibration encoding strategy. Experimental results demonstrate that the multimodal approach significantly outperforms unimodal conditions, achieving a spatial accuracy of 5.8 mm, a system usability score of 88.1, and notable reductions in cognitive load while enhancing user confidence during task execution.
Blind users face significant accessibility barriers in VR due to difficulties perceiving spatial information—such as direction, distance, and motion—of virtual objects. Method: This study presents the first systematic comparison, conducted with blind participants, of two tactile feedback modalities—dorsal-hand vibration versus skin stretch—for spatial information conveyance. We developed a custom dorsal-hand haptic device, a VR spatial rendering engine, and a dual-modality actuation control system, and evaluated performance using standardized experimental protocols with 10 blind participants. Contribution/Results: Skin stretch feedback significantly outperformed vibration: spatial position identification accuracy improved by 37%, and motion trajectory discrimination accuracy increased by 42%. Based on these findings, we propose evidence-based design guidelines for skin stretch haptics tailored to accessible VR. This work establishes a novel paradigm for high-fidelity spatial haptic interaction and provides empirical foundations for inclusive VR interface design.
Optical see-through augmented reality (OST-AR) imposes high visual workload during manual tasks and lacks reliable depth cues under occlusion or low-light conditions. To address these limitations, we propose a multimodal guidance approach integrating OST-AR with directional vibrotactile feedback delivered via a custom six-motor wristband. Our method introduces a novel “pull-based vibration” metaphor to jointly encode spatial direction and operational state. The system incorporates real-time handheld tool tracking, OST-AR visualization, and a millisecond-precise multimodal synchronization mechanism. User studies demonstrate significant improvements over unimodal baselines: 23.6% higher spatial guidance accuracy, 18.4% faster gesture completion time, enhanced cognitive efficiency, 94.2% tactile pattern recognition accuracy, and markedly increased user satisfaction—effectively overcoming the perceptual bottlenecks inherent in single-modality interfaces.
Current web-based 3D surface and point cloud visualization tools rely heavily on visual interaction, rendering them inaccessible to blind and low-vision (BLV) users in browser environments. To address this, we propose DIXTRAL—the first browser-native, synchronous multimodal 3D data visualization system designed specifically for BLV users. It integrates data sonification, dynamic textual descriptions, and optional visual feedback, and supports keyboard and game controller input. Its interaction logic was co-designed with BLV stakeholders and refined through iterative user studies. Experimental evaluation demonstrates that DIXTRAL significantly improves BLV users’ ability to recognize structural patterns in 3D scalar fields, perform spatial orientation, and conduct efficient exploratory analysis. This work contributes a reusable architectural paradigm and empirically grounded design guidelines for inclusive scientific visualization.
This work addresses the spatial misalignment challenge between haptic feedback and virtual objects in augmented reality (AR). We propose the first flexible, cross-platform architecture enabling deep integration of mid-air haptics (Ultrahaptics) with AR, supporting HoloLens, iOS, and diverse haptic devices—including wearable, grasp-based, and ultrasonic mid-air systems. Our approach leverages AR spatial registration, haptic-visual synchronized rendering, and cross-device semantic mapping to achieve high-fidelity haptic representation and spatial consistency of virtual objects. User studies demonstrate that mid-air haptics significantly improves shape recognition accuracy (+32%) and scaling task completion rate (+41%). We validate the architecture’s feasibility through two applications—Form Inspector and Simon Game—and uncover systematic user expectation mismatches regarding haptic metaphors (e.g., virtual buttons), thereby informing the evolution of haptic AR interface design principles.
This work addresses the significant accessibility challenges posed by the growing prevalence of social virtual reality (Social VR) for blind and low-vision users. We propose and evaluate an AI-powered voice navigation assistant driven by a large language model (LLM) to support navigation and social interaction within immersive VR environments. Through a user study involving 16 blind and low-vision participants, we reveal— for the first time—the assistant’s dual role as both a functional tool and a social companion in Social VR, alongside diverse user interaction strategies. Our findings demonstrate the effectiveness of LLM-based assistance in enhancing Social VR accessibility and offer critical design insights for future AI-augmented systems tailored to visually impaired users.
Current virtual reality systems struggle to support free-form exploration for blind and low-vision users, as they typically rely on visual feedback or are constrained by predefined menus and audio beacons. This work proposes a discovery-driven interaction paradigm that leverages natural head and hand movements, integrating multimodal feedback—including responsive spatial audio, directional haptics, and text-to-speech—to enable progressive discovery and navigation of environments and objects. A user study (N=12) demonstrates that the system significantly enhances exploratory freedom, with participants strongly preferring its discovery-based interaction; notably, task performance and cognitive load remained unaffected. This study presents the first virtual reality system to support truly free-form exploration for people who are blind or have low vision.
This study addresses the challenge users face in intuitively understanding and activating superhuman augmentation capabilities in virtual reality. Grounded in affordance theory, the research integrates user-centered design with expert participatory design, engaging professional designers to co-create avatars that effectively communicate both the presence and interaction modalities of such enhancements. The work presents the first systematic set of 16 design guidelines—comprising both general and category-specific principles—for representing augmented abilities in VR. Through controlled experiments and external evaluations, these guidelines were rated by users as clear and practical, and avatars designed following them demonstrated significantly improved intuitiveness. The framework has been successfully deployed across four distinct VR scenarios, offering a reusable design paradigm for augmented reality interactions.
This study addresses the lack of effective non-visual access to 3D data visualizations for blind and low-vision users in STEM domains. Through an empirically grounded co-design approach, the authors collaborated with accessibility experts across two iterative cycles, employing low-fidelity tactile probes and high-fidelity web prototypes to formulate a design protocol that translates tactile knowledge into digital interfaces. The resulting system innovatively integrates multimodal interaction techniques—including referential sonification, spatial and volumetric audio rendering, and configurable buffer aggregation—to significantly enhance the accuracy and learnability of non-visual 3D data analysis. User evaluations demonstrate that the tool effectively supports core analytical tasks such as directional orientation, peak identification, trend comparison, gradient tracing, and discovery of occluded features, offering practical design guidelines and a viable technical pathway toward accessible 3D visualization.
This study addresses the challenges faced by visually impaired individuals in learning physical movements—such as yoga or gymnastics—due to the lack of effective non-visual instructional tools. To bridge this gap, the authors introduce a novel approach that integrates high-fidelity 3D tactile human models with a user-centered participatory design process, resulting in custom 3D-printed models for blind learners. These models incorporate tactile markers to represent both static postures and continuous motion sequences. Findings from user studies demonstrate that, compared to conventional teaching methods, the proposed models significantly improve the speed and accuracy of movement comprehension, reduce learner uncertainty, and receive higher ratings in usability and motivation. The results indicate that the tactile models effectively enhance spatial awareness and facilitate more effective motor learning among visually impaired users.