Emotion-Conditioned Short-Horizon Human Pose Forecasting with a Lightweight Predictive World Model

📅 2026-04-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing short-term human pose forecasting methods, which predominantly rely on geometric motion cues while neglecting the influence of emotion on movement dynamics. The authors propose a lightweight autoregressive world model that integrates emotion embeddings—extracted from facial expressions—with pose keypoints through a learnable normalized gating mechanism, enabling 15-step roll-out predictions. Built upon a two-layer LSTM architecture, the model is evaluated on two emotion-annotated pose video datasets. Experimental results demonstrate that the normalized gating mechanism substantially improves prediction accuracy for emotion-driven motion sequences. Furthermore, counterfactual perturbation analyses reveal that predicted trajectories are sensitive to emotional inputs, confirming that emotion embeddings serve as effective conditioning signals rather than redundant features.

Technology Category

Computer Vision: Biometrics, Face, Gesture & PoseHumans and AI: Human-Aware Planning and Behavior PredictionReasoning under Uncertainty: Relational Probabilistic Models

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Short-term human pose prediction plays a crucial role in interactive systems, assistive robots, and emotion-aware human-computer interaction[1-3]. While current trajectory prediction models primarily rely on geometric motion cues, they often overlook the underlying emotional signals influencing human motion dynamics[4-5]. This paper investigates whether facial expression-derived emotion embeddings can provide auxiliary conditional signals for short-term pose prediction. To further evaluate multimodal conditionation in a recursive prediction setting, we propose a lightweight autoregressive predictive world model that performs 15-step rolling pose prediction. This framework combines pose keypoints with emotion embeddings through a learnable gating mechanism and performs autoregressive unfolding prediction using a recurrent sequence model based on a two-layer LSTM architecture. Experiments were conducted on two small-scale pose-emotion video datasets: controlled motion sequences with minimal facial expression changes and, natural emotion-driven motion sequences with considerable facial expression changes. The results show that simple multimodal fusion does not consistently improve prediction accuracy, while normalized gating fusion significantly enhances the performance of emotion-driven motion sequences. Furthermore, counterfactual perturbation experiments demonstrate that the predicted trajectory exhibits measurable sensitivity to changes in multimodal input, suggesting that facial expression embeddings act as auxiliary conditional signals rather than redundant features. In summary, these results indicate that incorporating facial expression-derived emotion embeddings into emotion-conditional short-term pose prediction based on a lightweight predictive world model architecture is a feasible approach.
Problem

Research questions and friction points this paper is trying to address.

human pose forecasting
emotion conditioning
facial expression
multimodal fusion
short-horizon prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

emotion-conditioned pose forecasting
lightweight predictive world model
multimodal gating fusion
facial expression embedding
autoregressive pose prediction
🔎 Similar Papers
No similar papers found.
J
Jingni Huang
P
Peter Bloodsworth