Score
Designs, trains, or analyzes compact, discrete representations of spatiotemporal motion (e.g., trajectory or gesture) patterns as reusable building blocks — including shape-aware and semantically anchored primitives — for use in planning, retrieval, decoding, or alignment. This work also covers mapping motion features to template slots and generating standardized, template-based natural-language verbalizations or semantic anchors that support retrieval, supervision, and downstream processing.
This paper presents a systematic review of movement primitive approaches in robot control, with a focus on learning from human demonstrations to generate complex action sequences. Integrating chronological and systematic perspectives, it comprehensively traces the theoretical evolution of movement primitives, key technical advances—including spring-damper modeling, probabilistic coupling of multiple demonstration trajectories, and neural network applications in high-dimensional systems—and their empirical effectiveness in tasks such as grasping and throwing. The study offers an in-depth comparative analysis of prevailing frameworks, establishes for the first time a structured developmental trajectory of the field, and clearly identifies current open challenges and practical limitations, thereby providing both theoretical guidance and a practical roadmap for research in robotic motor skill learning.
This work addresses the limitation of existing speech-gesture co-modeling approaches, which often fail to capture the communicative intent of semantic gestures and remain confined to low-level motion features. To overcome this, the authors propose a Semantic Motion Anchor mechanism that discretizes 3D gestures into body-hand action primitives and converts them into structured natural language descriptions aligned with spoken text, thereby establishing a semantics-driven cross-modal contrastive learning framework. Notably, this approach pioneers the use of natural language—encoding both physical form and communicative intent—as gesture anchors, moving beyond conventional end-to-end continuous embedding alignment. Evaluated on the BEAT2 dataset, the method achieves an 8.2% improvement in text-to-gesture retrieval R@1, outperforms state-of-the-art methods in bidirectional retrieval, and generates outputs significantly preferred by users, demonstrating the efficacy of semantic anchors in conveying communicative intent.
This work addresses the challenge that existing motion understanding methods struggle to enable large language models (LLMs) to perform fine-grained reasoning about complex actions due to a lack of geometric alignment between motion quantization and semantic embeddings. To bridge this gap, the authors propose a geometry-aware motion–language joint modeling framework that introduces, for the first time, an orthogonality constraint between the motion codebook and the LLM embedding space. A two-stage regularization strategy is designed to balance geometric consistency with semantic adaptability. The approach integrates a Gumbel-Softmax differentiable discrete decoder, sparse orthogonal projection mapping, and an orthogonality-regularized training mechanism. Evaluated on HumanML3D, the method achieves a 20% improvement over the current state of the art, demonstrating the effectiveness of geometric alignment in enhancing LLMs’ motion reasoning capabilities.
This study investigates the internal mechanisms underlying spatiotemporal reasoning in vision-language models (VLMs), with a focus on how spatial and textual representations are integrated. Through causal interventions, linear probing, representational analysis, and cross-modal activation alignment, the work systematically demonstrates the widespread presence of linear spatial and temporal identifiers in both image and video VLMs. It further reveals, for the first time, that these mechanisms modulate belief states in intermediate model layers. This insight not only offers a novel perspective for improving VLM interpretability and alignment design but also serves as a diagnostic tool to uncover model limitations and generate informative training signals.
This work addresses the lack of interpretability in existing skeleton-based action recognition models, which typically operate as black boxes. The authors propose a concept-driven interpretable framework that reformulates action recognition as first-order logical reasoning grounded in motion primitives. Their approach employs a spatiotemporal skeleton encoder and a concept decoder to learn differentiable spatiotemporal motion concepts, which are instantiated as logical predicates. By integrating a large language model to align atomic action semantics, the method constructs a shared conceptual space bridging perception and reasoning. This is the first effort to incorporate differentiable first-order logic into skeleton-based action recognition, achieving competitive accuracy on the NTU RGB+D 60/120 and NW-UCLA benchmarks while generating human-readable logical rules that enable explicit, interpretable action understanding.
This work addresses the challenge of aligning natural language prompts with spatiotemporal event dynamics in videos. We propose a language-driven 4D dynamic scene understanding framework that, for the first time, embeds natural language into a differentiable, temporally extended 3D Gaussian Splatting representation, enabling end-to-end text-to-dynamic-3D spatiotemporal localization. Our method integrates a CLIP text encoder with a lightweight spatiotemporal feature distillation module to construct a semantically consistent and geometrically accurate 4D Gaussian field. Evaluated on human and animal 3D video datasets, it reduces spatiotemporal localization error by 37% over baseline methods and supports real-time interactive querying. Key contributions are: (1) the first joint modeling of 4D Gaussian Splatting and linguistic modalities; (2) overcoming static or single-frame semantic alignment limitations via cross-frame semantic-geometric co-optimization; and (3) introducing the first differentiable, efficient, and interactive paradigm for language-guided dynamic 3D localization.
为解决长期任务中动态物体状态转换的自然语言查询问题,提出Linguistic Trajectory Encoding方法,结合自然语言、稀疏空间锚点和视觉锚点压缩运动历史。
This study addresses the challenge that existing data exploration tools struggle to accurately interpret users’ analytical intent when expressed in unstructured forms within spatiotemporal datasets. To bridge this gap, the authors propose a multimodal query system integrating freehand sketching, natural language, and visual annotations. Central to their approach is the concept of “proxemic semantics,” which captures how users disambiguate references through the relative spatial arrangement of multimodal elements within a unified interaction space. The system employs a hybrid architecture combining geometric sketch matching with vision-language models (VLMs), enabling joint pattern matching and semantic constraint-based query parsing. A user study with 20 participants empirically validates the stability of proxemic semantics, offering both empirical grounding and design implications for multimodal data exploration interfaces.
研究通过分析V-JEPA 2和VideoMAE-v2模型,探讨了视频基础模型中时空表示的编码内容、出现位置及几何组织方式,并使用轻量级探针来发现三种时间属性。
This study systematically evaluates the representational capabilities of vision-language models (VLMs) and video generation models (VGMs) on spatial intelligence tasks. Using a frozen-feature probing approach, the analysis compares their performance across three dimensions: semantic labeling, instance grouping, and 3D geometric prediction. The work reveals, for the first time, a complementary relationship between VLMs and VGMs in spatial understanding: VLMs excel at semantic and instance-level recognition, whereas VGMs demonstrate superior modeling of geometric structure and camera motion dynamics. Notably, a simple fusion of their representations yields substantial gains in overall performance, simultaneously enhancing both semantic accuracy and geometric fidelity.
研究通过空间音频和文本描述生成人体动作的问题,提出MoSAT框架,利用分层交叉注意力机制生成符合意图和环境线索的动作。