🤖 AI Summary
This study addresses the challenges of overlapping multiple actions, difficult limb dependency modeling, and high computational overhead in text-to-motion retrieval by proposing a Hierarchical Multi-Stream Motion-Aware Network. The method introduces a novel torso-centric interaction mechanism within a three-stream architecture—comprising upper limbs, lower limbs, and torso—to disentangle body-part motions while explicitly capturing inter-limb interactions. By integrating a tailored torso attention module with lightweight feature fusion, the framework efficiently processes long-sequence compound descriptions. Experimental results demonstrate that this approach significantly improves retrieval accuracy for both simple and complex multi-action queries, enhancing fine-grained relational understanding while maintaining low computational complexity. Ultimately, the proposed framework achieves an effective balance between efficiency and interpretability in motion retrieval tasks.
📝 Abstract
Accurate retrieval of human motions is a crucial first step in text-guided human motion modeling and synthesis, as it selects semantically relevant sequences from large datasets and provides grounded references for downstream tasks. Retrieving motions from natural language descriptions remains challenging because sentences can describe multiple actions, overlapping movements, and intricate dependencies between body parts. Existing methods often focus on simple, single-action descriptions and typically process body parts independently or by merely concatenating features, without explicitly modeling how torso movements influence other parts. In addition, their processing pipelines often rely on computationally heavy models, introducing considerable overhead, particularly when modeling longer or more complex motion sequences. This limits learning discriminative motion-pattern representations, reducing retrieval accuracy, interpretability, and efficiency in practical applications. To address these limitations, we propose HUMAN-TCI, a Hierarchical Multi-Stream Motion-Aware Network for text-guided human motion retrieval. HUMAN-TCI employs a three-stream architecture that separately models upper-body, lower-body, and torso motions while explicitly capturing their interactions, allowing torso-related movements to influence the positioning and dynamics of other body parts. By incorporating tailored torso attention, our model effectively recognizes complex human motion patterns, captures fine-grained motion relationships and handles complex multi-action descriptions. Our framework supports retrieval for both simple, single-action sentences and long, compositional descriptions containing sequential or overlapping actions without relying on complex models.