Comparing Utility of Inertial, Occupancy, Semantic, and Intent Information in Human Motion Prediction During Daily Tasks

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited accuracy of robot pedestrian motion prediction in complex indoor environments. Leveraging data collected via Meta Aria glasses, we propose a human motion diffusion model that integrates multi-source information, including inertial measurements, occupancy maps, semantic context, and eye-gaze intention cues. This work is the first to quantify the individual contributions of these information sources in ambiguous scenarios, revealing the critical role of explicit intention and gaze data in predicting deceleration behaviors. Experimental results demonstrate that the proposed approach improves prediction accuracy by 42% over a constant-velocity baseline, confirming that multimodal fusion significantly reduces prediction errors.
📝 Abstract
Accurate human motion prediction is crucial for robotic systems operating around people, particularly in complex indoor spaces. In this study, we assess the relative importance of different sources of information in indoor motion prediction with a human motion diffusion model. We collected a dataset of nine naive human subjects conducting simulated indoor daily activities while wearing a pair of Meta Aria glasses. This dataset includes ten buildings from a university campus, and encompasses 238 minutes of navigation between daily tasks. Overall, we demonstrate a 42% improvement beyond a constant velocity baseline. Including body motion, scene representation, and eye gaze fixation data significantly reduced prediction error. Semantic information was found to be useful for indoor motion prediction, but to a lesser degree than in outdoor navigation. Providing explicit intent information reduced error beyond any other addition, suggesting that incorporating explicit intent estimation or user input are fundamental for finer prediction of indoor motion. One of the few indicators of intent, eye gaze fixation, was found to be especially useful in predicting deceleration, and provided basic spatial information to the model in the absence of an occupancy map. These results are a first step towards predicting human motion in highly ambiguous indoor scenarios. Code will be made public upon acceptance. Project page: https://human-motion-diffusion.github.io/
Problem

Research questions and friction points this paper is trying to address.

Human Motion Prediction
Indoor Navigation
Intent Information
Semantic Information
Eye Gaze Fixation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human Motion Prediction
Diffusion Model
Intent Estimation
Eye Gaze Fixation
Indoor Navigation
💼 Related Jobs
No related jobs found.
M
Max Burns
Department of Mechanical Engineering, Stanford University, Stanford CA, USA
M
Maisha Khanum
Department of Mechanical Engineering, Stanford University, Stanford CA, USA
Monroe Kennedy III
Monroe Kennedy III
Assistant Professor of Mechanical Engineering, Stanford University
RoboticsRobotic AssistantsAssistive RoboticsRobotic ManipulationDynamics and Controls
S
Steven H. Collins
Department of Mechanical Engineering, Stanford University, Stanford CA, USA