Score
Combining and synchronizing multi-sensor data streams in real time to estimate state and update shared beliefs (e.g., pose, exposure), integrating modality-specific signals such as tactile feedback with learned predictions.
Tactile perception in embodied intelligence is hindered by spatial sparsity and the absence of global semantic context, while research on multimodal tactile fusion lacks a unified framework. This work systematically reviews relevant literature up to Q1 2026 and introduces, for the first time, a hierarchical taxonomy encompassing data modalities—such as tactile–visual and tactile–language—and three methodological pillars: perceptual recognition, cross-modal generation, and multimodal interaction. By integrating advances in deep learning and large language models, the study comprehensively surveys multimodal datasets, core algorithms, sensing hardware, and evaluation benchmarks, thereby clarifying the field’s developmental trajectory and offering a coherent theoretical foundation and systematic reference for future research.
Multimodal human activity recognition in natural scenes faces challenges in fusing asynchronous, heterogeneous sensor data (video, audio, RFID) and lacks interpretability. Method: This paper proposes a reproducible multimodal fusion framework that jointly addresses temporal alignment, feature standardization, and semantic parsing; supports early, late, and hybrid fusion strategies; and incorporates sparse RFID signals to enhance discriminative capability. Interpretability is achieved via t-SNE visualization, attention heatmaps, and waveform superposition for modality-wise contribution analysis. Contribution/Results: The work presents the first systematic evaluation of multimodal synergy in temporal consistency and discriminability. Experiments show late fusion achieves optimal performance; integrating RFID improves classification accuracy by over 50% and significantly boosts macro-averaged ROC-AUC, demonstrating the framework’s dual advantages in accuracy and interpretability.
This work addresses the challenge that monolithic large language models struggle to balance accuracy and robustness when fusing heterogeneous multimodal sensor data, primarily due to prior biases and vulnerability to missing modalities. To overcome this limitation, the authors propose ConSensus, a novel training-free, single-round multi-agent collaboration framework. ConSensus decomposes perception tasks into modality-specific agents and integrates their outputs through a hybrid fusion strategy combining semantic aggregation and statistical consensus. This approach preserves cross-modal contextual understanding while substantially reducing computational overhead. Evaluated on five standard multimodal perception benchmarks, ConSensus achieves an average accuracy improvement of 7.1% over existing methods, matching the performance of iterative multi-agent debate approaches while reducing fusion token consumption by a factor of 12.7.
Scalable, high-resolution, multimodal synchronous acquisition tools are lacking in educational settings, hindering the practical deployment of learning analytics. To address this, we introduce Watch-DMLT and ViSeDOPS—two integrated systems enabling, for the first time, real-time, multi-user physiological (heart rate) and motion sensing via Fitbit Sense 2 smartwatches, synchronized with eye-tracking, video, and contextual annotations at millisecond-level temporal precision, supported by interactive visualization. The system was end-to-end deployed in classroom oral presentation tasks involving 65 students, demonstrating feasibility and efficacy for fine-grained, scalable learning analytics in authentic educational environments. Key contributions include: (1) the first classroom-scale, non-intrusive, multi-user multimodal synchronization framework achieving high temporal fidelity; and (2) an open-source toolchain and web-based visualization dashboard that substantially lowers technical barriers for multimodal educational research.
The Platform for Situated Intelligence (Psi) lacks compatibility with the Unity/Mixed Reality Toolkit (MRTK) ecosystem, severely hindering reproducibility in multimodal XR experiments. Method: This work introduces the first deep integration of Psi’s high-precision temporal alignment capabilities into Unity 2022.3 and MRTK3—bypassing StereoKit limitations—to enable bidirectional, real-time streaming of HoloLens 2 sensor data (AHAT, long-range depth, IMU, eye tracking, hand gestures, and audio). Built in C#, it leverages the Psi .NET SDK and incorporates microsecond-level time coordination, native serialization, and structured logging. Contribution/Results: The platform achieves sub-millisecond synchronization between head-mounted displays and immersive applications, significantly enhancing measurement fidelity and experimental reproducibility in human–robot interaction (HRI), human–computer interaction (HCI), and embodied intelligence research. The implementation is open-source.
Current artificial intelligence remains largely confined to digital modalities such as text, vision, and audio, falling short of comprehensively perceiving the rich multisensory world inhabited by humans. This work proposes a ten-year research vision for multisensory intelligence, introducing the first systematic framework that unifies AI with the full spectrum of human senses. Centered on three core directions—sensory perception, modeling, and human-AI collaboration—the framework integrates physiological, tactile, and environmental signals to advance key technologies including cross-modal alignment, unified representation learning, multisensory generation, and transfer learning, thereby overcoming fundamental bottlenecks in heterogeneous modality fusion. The contribution includes a comprehensive research roadmap accompanied by a suite of projects, open-source resources, and demonstration systems developed by the MIT Media Lab team, aiming to catalyze a new paradigm in multimodal human-AI interaction.
Existing systems struggle to support strict audio-visual synchronization, limiting the analysis of fine-grained temporal features in dialogue such as turn-taking, overlapping speech, and prosody. To address this challenge, this work proposes an end-to-end multimodal acquisition and calibration framework that treats synchronized audio and video as equally central modalities for the first time. By integrating a multi-camera array with multi-channel microphones under a unified temporal architecture, the system enables scalable, reproducible, high-quality recording. Standardized calibration and quality control procedures ensure high temporal consistency across modalities, yielding data that effectively supports fine-grained analysis of conversational behavior and data-driven modeling.
Existing vision-language-action models struggle to effectively integrate heterogeneous physical feedback signals, limiting their perception and decision-making capabilities in real-world settings. This work proposes MoSS, a framework that employs decoupled, modular sensory streams to flexibly incorporate multimodal physical signals such as tactile and torque data, and introduces a joint cross-modal self-attention mechanism to enhance modeling of contact dynamics. By combining a two-stage training strategy with an auxiliary task of future signal prediction, MoSS achieves stable and efficient fusion of multiple physical modalities. Real-robot experiments demonstrate that the proposed approach significantly improves action prediction performance, validating the efficacy and added value of synergistically integrating multimodal physical feedback.
This work addresses the challenge that certain modalities may introduce interference under specific inputs in multimodal fusion, thereby degrading model performance. To mitigate this issue, the authors propose a plug-and-play pre-fusion calibration module that leverages cross-modal summary contrast to extract supportive and conflicting cues, generating instance-level and dimension-level modulation signals. These signals dynamically enhance beneficial features while suppressing misleading information. The method achieves, for the first time, fine-grained, conflict-aware modulation prior to fusion and is compatible with both sequential and convolutional architectures. It consistently improves performance across five benchmark tasks—including emotion understanding and action recognition—and demonstrates enhanced robustness and consistency under modality missingness and data perturbations.
This work addresses the lack of a standardized, reusable framework for integrating physiological signals within the current ROS 2 ecosystem, which hinders effective assessment of users’ psychological states in human-robot interaction. To bridge this gap, we propose and implement a modular physiological data processing framework tailored for ROS 2, featuring the first extensible, standardized interfaces that support multi-source physiological sensor integration, high-precision time synchronization, context-aware logging, and multimodal data fusion. The framework enables real-time inference of user state indicators, significantly enhancing the reliability, traceability, and interoperability of physiological data analysis in ROS 2–based human-robot interaction systems.
This work addresses the rapid spread of misinformation in everyday conversations, where users struggle to verify claims in real time. The authors propose a wearable system that continuously monitors ambient audio to detect verifiable statements and integrates real-time web-based fact-checking to deliver immediate, subtle feedback via haptic nudges and glanceable visual cues. In the first empirical study of its kind (N=34), this body-integrated design significantly improved participants’ real-time accuracy in distinguishing true from false claims and encouraged proactive verification behaviors. The findings also uncover a nuanced tension between user trust in the system and the risk of overreliance, offering critical insights for the design of trustworthy AI-augmented assistance in human communication.