M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production

📅 2026-03-24
📈 Citations: 0
✨ Influential: 0
📄 PDF

Technology Category

Natural Language Processing: GenerationMachine Learning: Multimodal LearningComputer Vision: Multi-modal Vision

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Sign language production requires more than hand motion generation. Non-manual features, including mouthings, eyebrow raises, gaze, and head movements, are grammatically obligatory and cannot be recovered from manual articulators alone. Existing 3D production systems face two barriers to integrating them: the standard body model provides a facial space too low-dimensional to encode these articulations, and when richer representations are adopted, standard discrete tokenization suffers from codebook collapse, leaving most of the expression space unreachable. We propose SMPL-FX, which couples FLAME's rich expression space with the SMPL-X body, and tokenize the resulting representation with modality-specific Finite Scalar Quantization VAEs for body, hands, and face. M3T is an autoregressive transformer trained on this multi-modal motion vocabulary, with an auxiliary translation objective that encourages semantically grounded embeddings. Across three standard benchmarks (How2Sign, CSL-Daily, Phoenix14T) M3T achieves state-of-the-art sign language production quality, and on NMFs-CSL, where signs are distinguishable only by non-manual features, reaches 58.3% accuracy against 49.0% for the strongest comparable pose baseline.
Problem

Research questions and friction points this paper is trying to address.

sign language production
non-manual features
discrete tokenization
codebook collapse
3D motion representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-modal motion tokens
Non-manual features
Discrete tokenization
SMPL-FX
Sign language production