learn motion primitives

Designs, trains, or analyzes compact, discrete representations of spatiotemporal motion (e.g., trajectory or gesture) patterns as reusable building blocks — including shape-aware and semantically anchored primitives — for use in planning, retrieval, decoding, or alignment. This work also covers mapping motion features to template slots and generating standardized, template-based natural-language verbalizations or semantic anchors that support retrieval, supervision, and downstream processing.

learnmotionprimitives

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitation of existing speech-gesture co-modeling approaches, which often fail to capture the communicative intent of semantic gestures and remain confined to low-level motion features. To overcome this, the authors propose a Semantic Motion Anchor mechanism that discretizes 3D gestures into body-hand action primitives and converts them into structured natural language descriptions aligned with spoken text, thereby establishing a semantics-driven cross-modal contrastive learning framework. Notably, this approach pioneers the use of natural language—encoding both physical form and communicative intent—as gesture anchors, moving beyond conventional end-to-end continuous embedding alignment. Evaluated on the BEAT2 dataset, the method achieves an 8.2% improvement in text-to-gesture retrieval R@1, outperforms state-of-the-art methods in bidirectional retrieval, and generates outputs significantly preferred by users, demonstrating the efficacy of semantic anchors in conveying communicative intent.

co-speech gesturescommunicative intentgesture retrieval

This work addresses the challenge that existing motion understanding methods struggle to enable large language models (LLMs) to perform fine-grained reasoning about complex actions due to a lack of geometric alignment between motion quantization and semantic embeddings. To bridge this gap, the authors propose a geometry-aware motion–language joint modeling framework that introduces, for the first time, an orthogonality constraint between the motion codebook and the LLM embedding space. A two-stage regularization strategy is designed to balance geometric consistency with semantic adaptability. The approach integrates a Gumbel-Softmax differentiable discrete decoder, sparse orthogonal projection mapping, and an orthogonality-regularized training mechanism. Evaluated on HumanML3D, the method achieves a 20% improvement over the current state of the art, demonstrating the effectiveness of geometric alignment in enhancing LLMs’ motion reasoning capabilities.

embedding spacegeometric alignmentlarge language models

This study investigates the internal mechanisms underlying spatiotemporal reasoning in vision-language models (VLMs), with a focus on how spatial and textual representations are integrated. Through causal interventions, linear probing, representational analysis, and cross-modal activation alignment, the work systematically demonstrates the widespread presence of linear spatial and temporal identifiers in both image and video VLMs. It further reveals, for the first time, that these mechanisms modulate belief states in intermediate model layers. This insight not only offers a novel perspective for improving VLM interpretability and alignment design but also serves as a diagnostic tool to uncover model limitations and generate informative training signals.

mechanism interpretabilityspatial representationspatiotemporal reasoning

This work addresses the lack of interpretability in existing skeleton-based action recognition models, which typically operate as black boxes. The authors propose a concept-driven interpretable framework that reformulates action recognition as first-order logical reasoning grounded in motion primitives. Their approach employs a spatiotemporal skeleton encoder and a concept decoder to learn differentiable spatiotemporal motion concepts, which are instantiated as logical predicates. By integrating a large language model to align atomic action semantics, the method constructs a shared conceptual space bridging perception and reasoning. This is the first effort to incorporate differentiable first-order logic into skeleton-based action recognition, achieving competitive accuracy on the NTU RGB+D 60/120 and NW-UCLA benchmarks while generating human-readable logical rules that enable explicit, interpretable action understanding.

interpretabilitylogical reasoningmotion primitives

4-LEGS: 4D Language Embedded Gaussian Splatting

Oct 14, 2024
GF
Gal Fiebelman
🏛️ Tel Aviv University | Google Research

This work addresses the challenge of aligning natural language prompts with spatiotemporal event dynamics in videos. We propose a language-driven 4D dynamic scene understanding framework that, for the first time, embeds natural language into a differentiable, temporally extended 3D Gaussian Splatting representation, enabling end-to-end text-to-dynamic-3D spatiotemporal localization. Our method integrates a CLIP text encoder with a lightweight spatiotemporal feature distillation module to construct a semantically consistent and geometrically accurate 4D Gaussian field. Evaluated on human and animal 3D video datasets, it reduces spatiotemporal localization error by 37% over baseline methods and supports real-time interactive querying. Key contributions are: (1) the first joint modeling of 4D Gaussian Splatting and linguistic modalities; (2) overcoming static or single-frame semantic alignment limitations via cross-frame semantic-geometric co-optimization; and (3) introducing the first differentiable, efficient, and interactive paradigm for language-guided dynamic 3D localization.

Connect language with dynamic 3D modelingEnable text prompts to localize video eventsLift spatio-temporal features to 4D representation

Latest Papers

What's happening recently
View more

This study addresses the challenge that existing data exploration tools struggle to accurately interpret users’ analytical intent when expressed in unstructured forms within spatiotemporal datasets. To bridge this gap, the authors propose a multimodal query system integrating freehand sketching, natural language, and visual annotations. Central to their approach is the concept of “proxemic semantics,” which captures how users disambiguate references through the relative spatial arrangement of multimodal elements within a unified interaction space. The system employs a hybrid architecture combining geometric sketch matching with vision-language models (VLMs), enabling joint pattern matching and semantic constraint-based query parsing. A user study with 20 participants empirically validates the stability of proxemic semantics, offering both empirical grounding and design implications for multimodal data exploration interfaces.

analytical intentdeictic disambiguationmultimodal data exploration

This study systematically evaluates the representational capabilities of vision-language models (VLMs) and video generation models (VGMs) on spatial intelligence tasks. Using a frozen-feature probing approach, the analysis compares their performance across three dimensions: semantic labeling, instance grouping, and 3D geometric prediction. The work reveals, for the first time, a complementary relationship between VLMs and VGMs in spatial understanding: VLMs excel at semantic and instance-level recognition, whereas VGMs demonstrate superior modeling of geometric structure and camera motion dynamics. Notably, a simple fusion of their representations yields substantial gains in overall performance, simultaneously enhancing both semantic accuracy and geometric fidelity.

pretraining paradigmspatial intelligenceVideo Generation Models

Hot Scholars

RW

Ruihai Wu

Peking University
computer visionrobotics
MS

Matteo Saveriano

Associate Professor, University of Trento
RoboticsMachine LearningAI
KO

Keisuke Okumura

University of Cambridge & National Institute of Advanced Industrial Science and Technology (AIST)
Multi-Agent Path PlanningMulti-Robot Coordination
SB

Stefano Baraldo

SUPSI - University of Applied Sciences of Southern Switzerland
Industrial RoboticsMetal Additive ManufacturingMachine Learning