Score
Designs and implements motion encoding systems that decompose temporal pose or motion sequences into anatomy-aligned, per-frame part tokens by training or constructing part-specific vector-quantized (VQ) codebooks and tokenization pipelines; these representations expose frame–part tokens for transformer-style attention, enable localized per-part control addressing, and compress motion into attention-addressable discrete codes. Analyzes and tunes choices such as part partitioning, codebook size, temporal consistency, and reconstruction fidelity to balance compression, controllability, and downstream attention-based modeling.
Existing discrete pose encoding methods struggle to model fine-grained motion details, resulting in limited expressiveness and poor disentanglement. To address this, we propose a hierarchical pose representation framework that integrates discrete pose codes with continuous residual features via Residual Vector Quantization (RVQ), significantly enhancing motion detail fidelity while preserving interpretability and controllability. Our method employs an autoregressive decoder for end-to-end text-to-motion generation and is trained on the HumanML3D dataset. Experiments demonstrate substantial improvements: the Fréchet Inception Distance (FID) drops from 0.041 to 0.015, and Top-1 R-Precision rises to 0.510. Qualitative evaluation confirms high precision, strong controllability, and semantic consistency in motion editing tasks. The core contribution lies in the first application of RVQ to text-driven motion generation, achieving an organic unification of discrete controllability and continuous expressiveness.
Existing approaches to human motion generation face two key bottlenecks: difficulty in modeling multi-scale motion patterns and limited compositional flexibility of discrete representations. This paper introduces MSQ—the first spatiotemporal multi-scale motion quantization framework—which extracts part-level spatial features via multiple encoders and jointly applies temporal interpolation and vector quantization to compress continuous motion into multi-granularity discrete tokens. Its core innovation lies in enabling zero-shot, fine-tuned-free token stitching and recombination across scales, significantly enhancing generative flexibility and cross-task generalization. MSQ integrates generative masked modeling to establish an end-to-end discrete representation learning paradigm. Extensive experiments demonstrate that MSQ outperforms state-of-the-art methods on diverse benchmarks for motion generation, editing, and conditional control tasks. Both quantitative metrics and qualitative analyses consistently validate its superiority.
This work addresses a critical limitation in existing text-to-motion generation methods, which treat motion codebook indices as unordered categorical tokens and thereby ignore their intrinsic kinematic geometric structure. To overcome this, the authors propose MoGeFlow, the first approach to explicitly reveal and leverage the measurable, non-random, and decoder-causal local geometric properties inherent in motion codebooks. By reformulating discrete token prediction as a generative task in a continuous geometric space, MoGeFlow operates on the PartVQ codebook through text-conditioned continuous normalizing flows, geometry-aware embedding generation, and codebook projection—all while keeping the motion decoder frozen to ensure structural consistency. Experiments demonstrate that MoGeFlow sets new state-of-the-art results on HumanML3D and KIT-ML, achieving record R-Precision scores, superior multi-modal distance, and the best FID, while also leading across multiple metrics on the MotionMillion benchmark.
Existing methods for humanoid motion generation typically employ a single codebook to uniformly quantize both low-frequency pose semantics and high-frequency physical dynamics, which struggles to adequately capture fine-grained motion details. To address this limitation, this work proposes a Dual-Stream Frequency-domain Tokenizer (DSFT), introducing for the first time a frequency-aware decoupling mechanism: base poses are extracted via DCT truncation, while physical dynamics are compressed using Byte Pair Encoding (BPE), yielding two separate token streams. Built upon the Qwen-3.5 architecture, the MotionVLA model autoregressively predicts base tokens followed by physical tokens. Experiments demonstrate that this approach reduces the diversity gap by over 50% on HumanML3D and improves action-condition consistency by 3.8% on MBench, validating the efficacy of frequency-domain decoupling even within a lightweight 2B-parameter model.
Extreme token compression (e.g., >99.9% reduction) in video large language models (VLMs) severely distorts spatiotemporal modeling, degrading long-video understanding. Method: We propose a dynamic video token representation framework: (1) formally define the novel task of *extreme short-token compression*; (2) decouple visual content from grid-level motion to construct a compact token backbone and a token dynamics graph; (3) introduce cross-dynamics attention to fuse motion semantics without increasing token count. Our method integrates token decoupling, object-level clustering, grid-motion representation, and adaptive/fixed-length compression subtasks. Results: The framework reduces theoretical complexity by compressing tokens to just 0.07% of the original number, while incurring only a 1.13% performance drop on downstream tasks and achieving substantial throughput gains—enabling efficient, high-fidelity long-video–language understanding.
This work addresses the high inference latency of Vision-Language-Action (VLA) models on GPUs, which hinders real-time deployment, and the inadequacy of existing acceleration methods in fully eliminating weight and computational redundancy. To overcome these limitations, the paper proposes VQVLA, a novel framework that introduces motion-aware dynamic quantization (MotionVQ) and codebook-index-based vectorized GEMM. By leveraging spatiotemporal centroid reuse and dynamic precision selection within an algorithm-hardware co-design paradigm, VQVLA substantially reduces memory traffic and redundant computation. Experimental results demonstrate that VQVLA achieves a 6.5× speedup over the A100 GPU and a 2.8× speedup over Dadu-Corki while maintaining task success rates and incurring negligible accuracy loss.