Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of real-time modeling of human-centric upper-body interactions with objects, simultaneously generating coherent full-body, hand, and facial motions while enabling controllable synthesis of discrete interaction states such as “grasping.” To this end, the authors propose a continuous–discrete joint control architecture that employs a multi-scale implicit motion representation to unify whole-body motion encoding without explicit retargeting, leverages natural language instructions to precisely guide discrete interaction states, incorporates a dedicated rendering pipeline to provide interaction-aware supervision signals, and utilizes model distillation to enable real-time streaming inference. The resulting system achieves 25 FPS on dual H100 GPUs, significantly improving motion fidelity, hand–object coordination, and fine-grained controllability of local interactions.
📝 Abstract
We present a real-time human-centric world model for upper-body interactive generation, aiming to synthesize coherent local world dynamics centered on a person, where coordinated body, hand, and facial motions evolve jointly with controllable human-object discrete interaction. To this end, we adopt a continuous-discrete joint control scheme with two complementary components: a continuous human state and a discrete interaction state. For continuous human-state control, we introduce a unified implicit representation based on multi-scale motion encoding, in which motion latents from the upper body, hands, and face are fused into a shared latent space. This multi-scale design improves expressiveness across different spatial scales, captures fine-grained human dynamics more effectively, and enables direct control without explicit retargeting. For discrete object interaction-state control, we represent object contact using a small set of language-encoded discrete interaction states, where text serves as an explicit interaction-state command, such as \emph{no contact} or \emph{grasp}, rather than an open-ended generation prompt, and we further construct a dedicated rendering pipeline for human-object interaction data to supervise such discrete interaction states. By combining continuous implicit human-state control with discrete interaction-state control, our model enables precise modeling of how a person moves and interacts with the local environment, including controllable changes to nearby scene states. Finally, we distill the model for efficient streaming real-time inference, achieving 25 FPS on two H100 GPUs. Experiments demonstrate improved fine-grained motion fidelity, more realistic hand-object coordination, and effective real-time interaction, establishing a practical step beyond motion reproduction toward real-time human-centric world modeling.
Problem

Research questions and friction points this paper is trying to address.

human-object interaction
real-time modeling
upper-body motion
world modeling
discrete interaction states
Innovation

Methods, ideas, or system contributions that make the work stand out.

human-centric world modeling
continuous-discrete control
multi-scale motion encoding
discrete interaction states
real-time inference