sensor fusion

Combining and synchronizing multi-sensor data streams in real time to estimate state and update shared beliefs (e.g., pose, exposure), integrating modality-specific signals such as tactile feedback with learned predictions.

sensorfusion

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Multimodal human activity recognition in natural scenes faces challenges in fusing asynchronous, heterogeneous sensor data (video, audio, RFID) and lacks interpretability. Method: This paper proposes a reproducible multimodal fusion framework that jointly addresses temporal alignment, feature standardization, and semantic parsing; supports early, late, and hybrid fusion strategies; and incorporates sparse RFID signals to enhance discriminative capability. Interpretability is achieved via t-SNE visualization, attention heatmaps, and waveform superposition for modality-wise contribution analysis. Contribution/Results: The work presents the first systematic evaluation of multimodal synergy in temporal consistency and discriminability. Experiments show late fusion achieves optimal performance; integrating RFID improves classification accuracy by over 50% and significantly boosts macro-averaged ROC-AUC, demonstrating the framework’s dual advantages in accuracy and interpretability.

Develops reproducible multimodal fusion framework for human activity recognitionEvaluates sensor modality contributions and interpretability in activity classificationTransforms raw asynchronous sensor data into aligned meaningful representations

This work addresses the challenge that monolithic large language models struggle to balance accuracy and robustness when fusing heterogeneous multimodal sensor data, primarily due to prior biases and vulnerability to missing modalities. To overcome this limitation, the authors propose ConSensus, a novel training-free, single-round multi-agent collaboration framework. ConSensus decomposes perception tasks into modality-specific agents and integrates their outputs through a hybrid fusion strategy combining semantic aggregation and statistical consensus. This approach preserves cross-modal contextual understanding while substantially reducing computational overhead. Evaluated on five standard multimodal perception benchmarks, ConSensus achieves an average accuracy improvement of 7.1% over existing methods, matching the performance of iterative multi-agent debate approaches while reducing fusion token consumption by a factor of 12.7.

cross-modal reasoningheterogeneous sensor dataLLM bias

Real-Time Multimodal Data Collection Using Smartwatches and Its Visualization in Education

Dec 02, 2025
AB
Alvaro Becerra
🏛️ Universidad Autónoma de Madrid

Scalable, high-resolution, multimodal synchronous acquisition tools are lacking in educational settings, hindering the practical deployment of learning analytics. To address this, we introduce Watch-DMLT and ViSeDOPS—two integrated systems enabling, for the first time, real-time, multi-user physiological (heart rate) and motion sensing via Fitbit Sense 2 smartwatches, synchronized with eye-tracking, video, and contextual annotations at millisecond-level temporal precision, supported by interactive visualization. The system was end-to-end deployed in classroom oral presentation tasks involving 65 students, demonstrating feasibility and efficacy for fine-grained, scalable learning analytics in authentic educational environments. Key contributions include: (1) the first classroom-scale, non-intrusive, multi-user multimodal synchronization framework achieving high temporal fidelity; and (2) an open-source toolchain and web-based visualization dashboard that substantially lowers technical barriers for multimodal educational research.

Lack of scalable synchronized multimodal data tools in educationNeed for real-time monitoring of physiological and behavioral signalsVisualization of synchronized data for learning analytics in classrooms

psiUnity: A Platform for Multimodal Data-Driven XR

Nov 07, 2025
AA
Akhil Ajikumar
🏛️ Georgia Institute of Technology

The Platform for Situated Intelligence (Psi) lacks compatibility with the Unity/Mixed Reality Toolkit (MRTK) ecosystem, severely hindering reproducibility in multimodal XR experiments. Method: This work introduces the first deep integration of Psi’s high-precision temporal alignment capabilities into Unity 2022.3 and MRTK3—bypassing StereoKit limitations—to enable bidirectional, real-time streaming of HoloLens 2 sensor data (AHAT, long-range depth, IMU, eye tracking, hand gestures, and audio). Built in C#, it leverages the Psi .NET SDK and incorporates microsecond-level time coordination, native serialization, and structured logging. Contribution/Results: The platform achieves sub-millisecond synchronization between head-mounted displays and immersive applications, significantly enhancing measurement fidelity and experimental reproducibility in human–robot interaction (HRI), human–computer interaction (HCI), and embodied intelligence research. The implementation is open-source.

Bridging multimodal data management platform with Unity ecosystemEnabling real-time synchronized streaming of XR sensor dataExtending temporal coordination capabilities to HoloLens development environment

Current artificial intelligence remains largely confined to digital modalities such as text, vision, and audio, falling short of comprehensively perceiving the rich multisensory world inhabited by humans. This work proposes a ten-year research vision for multisensory intelligence, introducing the first systematic framework that unifies AI with the full spectrum of human senses. Centered on three core directions—sensory perception, modeling, and human-AI collaboration—the framework integrates physiological, tactile, and environmental signals to advance key technologies including cross-modal alignment, unified representation learning, multisensory generation, and transfer learning, thereby overcoming fundamental bottlenecks in heterogeneous modality fusion. The contribution includes a comprehensive research roadmap accompanied by a suite of projects, open-source resources, and demonstration systems developed by the MIT Media Lab team, aiming to catalyze a new paradigm in multimodal human-AI interaction.

artificial intelligencehuman-AI interactionmultimodal perception

Latest Papers

What's happening recently
View more

Existing systems struggle to support strict audio-visual synchronization, limiting the analysis of fine-grained temporal features in dialogue such as turn-taking, overlapping speech, and prosody. To address this challenge, this work proposes an end-to-end multimodal acquisition and calibration framework that treats synchronized audio and video as equally central modalities for the first time. By integrating a multi-camera array with multi-channel microphones under a unified temporal architecture, the system enables scalable, reproducible, high-quality recording. Standardized calibration and quality control procedures ensure high temporal consistency across modalities, yielding data that effectively supports fine-grained analysis of conversational behavior and data-driven modeling.

audio-visual synchronizationconversational interactionhuman motion recording

Existing vision-language-action models struggle to effectively integrate heterogeneous physical feedback signals, limiting their perception and decision-making capabilities in real-world settings. This work proposes MoSS, a framework that employs decoupled, modular sensory streams to flexibly incorporate multimodal physical signals such as tactile and torque data, and introduces a joint cross-modal self-attention mechanism to enhance modeling of contact dynamics. By combining a two-stage training strategy with an auxiliary task of future signal prediction, MoSS achieves stable and efficient fusion of multiple physical modalities. Real-robot experiments demonstrate that the proposed approach significantly improves action prediction performance, validating the efficacy and added value of synergistically integrating multimodal physical feedback.

action predictionheterogeneous sensory signalsmultimodal integration

This work addresses the challenge that certain modalities may introduce interference under specific inputs in multimodal fusion, thereby degrading model performance. To mitigate this issue, the authors propose a plug-and-play pre-fusion calibration module that leverages cross-modal summary contrast to extract supportive and conflicting cues, generating instance-level and dimension-level modulation signals. These signals dynamically enhance beneficial features while suppressing misleading information. The method achieves, for the first time, fine-grained, conflict-aware modulation prior to fusion and is compatible with both sequential and convolutional architectures. It consistently improves performance across five benchmark tasks—including emotion understanding and action recognition—and demonstrates enhanced robustness and consistency under modality missingness and data perturbations.

contextual adjustmentcross-modal conflictmodality calibration

This work addresses the lack of a standardized, reusable framework for integrating physiological signals within the current ROS 2 ecosystem, which hinders effective assessment of users’ psychological states in human-robot interaction. To bridge this gap, we propose and implement a modular physiological data processing framework tailored for ROS 2, featuring the first extensible, standardized interfaces that support multi-source physiological sensor integration, high-precision time synchronization, context-aware logging, and multimodal data fusion. The framework enables real-time inference of user state indicators, significantly enhancing the reliability, traceability, and interoperability of physiological data analysis in ROS 2–based human-robot interaction systems.

human-robot interactionmental state estimationphysiological signals

This work addresses the rapid spread of misinformation in everyday conversations, where users struggle to verify claims in real time. The authors propose a wearable system that continuously monitors ambient audio to detect verifiable statements and integrates real-time web-based fact-checking to deliver immediate, subtle feedback via haptic nudges and glanceable visual cues. In the first empirical study of its kind (N=34), this body-integrated design significantly improved participants’ real-time accuracy in distinguishing true from false claims and encouraged proactive verification behaviors. The findings also uncover a nuanced tension between user trust in the system and the risk of overreliance, offering critical insights for the design of trustworthy AI-augmented assistance in human communication.

fact-checkingmisinformationreal-time verification

Hot Scholars

IK

Itzik Klein

University of Haifa
RoboticsInertial SensingData-Driven NavigationAUV
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
LX

Lihua Xie

Professor of Electrical Engineering, Nanyang Technological University
Robust controlNetworked ControlMult-agent Systems
BG

Banglei Guan

National University of Defense Technology
PhotomechanicsVideometrics
MS

Martin Saska

Czech Technical University in Prague
roboticsautonomous systemsmulti-robot systemsUAV swarms