Score
Designs, builds, and evaluates models and training methods that produce compact continuous or structured representations (embeddings or neural implicit functions) for a wide range of modalities — e.g., text, images, audio, video, 3D geometry, code, molecules, and user behavior — that capture semantic, temporal, motion and content properties. This work encompasses architecting encoders/decoders and cross-encoders, learning alignment and fusion across modalities, creating temporal- or motion-aware and neural implicit scene representations, and measuring representation quality, transferability, and downstream utility.
This study systematically investigates video-text cross-modal representation alignment, focusing on modern encoders’ spatiotemporal modeling capabilities and their relationship with downstream performance. Method: We propose parameterized test-time scaling laws to quantitatively link semantic alignment degree with video understanding ability; design a novel temporal reasoning benchmark to overcome limitations of conventional zero-shot classification evaluation; and integrate multi-frame video encoding, text-set alignment, and regression-based modeling to jointly learn static and dynamic representations. Results: Experiments demonstrate that strong text alignment significantly enhances general-purpose video representation quality, and alignment metrics reliably predict model performance across diverse video understanding tasks. Our work advances the understanding of multimodal model internals and provides an interpretable pathway for alignment optimization.
Direct post-training alignment between pre-trained unimodal 3D encoders and text feature spaces yields limited performance due to semantic-geometric entanglement in 3D representations. Method: We propose projecting 3D features into a low-dimensional, semantically and geometrically disentangled subspace—jointly optimized via PCA and contrastive learning—to enable efficient post-training alignment without modifying the pre-trained 3D encoder or requiring multimodal joint training. Contribution/Results: This work establishes the first baseline for post-training alignment of unimodal 3D–text representations. We empirically reveal that semantic and geometric information in 3D features are naturally separable in such low-dimensional subspaces. Our method significantly improves cross-modal matching and retrieval accuracy (average +12.7%), validating the existence of implicit shared structural priors between 3D and text modalities. It offers a lightweight, plug-and-play paradigm for adapting pre-trained 3D foundation models to cross-modal tasks.
Existing video understanding research overlooks how structural characteristics of datasets—such as motion complexity, temporal span, hierarchical composition, and multimodal richness—guide the evolution of model architectures. Method: We propose a dataset-centric analytical framework that systematically interprets mainstream architectures—including two-stream networks, 3D CNNs, RNNs, Transformers, and multimodal foundation models—as responses to dataset-imposed inductive biases. Our approach integrates literature review with architecture–bias–task alignment analysis, unifying inductive bias theory and multimodal learning paradigms. Contribution/Results: We establish, for the first time, a unified “dataset → inductive bias → model design” framework, revealing the intrinsic logic underlying architectural evolution. The framework yields principled, generalizable design guidelines for video understanding models that balance scalability and task adaptability, advancing both theoretical understanding and practical model development.
This work addresses the suboptimal performance in multimodal representation alignment caused by modality gaps and data scarcity. To this end, the authors propose a disentangled representation learning framework based on shared and modality-specific codebooks. Leveraging a compositional vector quantization mechanism, the method decomposes multimodal features into shared semantic components and modality-unique components, and employs a progressive alignment strategy to optimize the alignment space without requiring fully paired data. The unified shared codebook effectively bridges the modality gap, while the modality-specific codebooks mitigate dominant-modality bias, enabling more balanced multimodal fusion. The approach achieves state-of-the-art performance across classification and retrieval tasks spanning nine modalities, including text, images, video, and audio.
This work addresses the challenge of constructing open-source large language models (LLMs) capable of processing arbitrary modalities. We propose OmniAlignNet, a novel architecture that integrates temporal embedding grouping with constrained rotational time embeddings to achieve precise cross-modal alignment and robust temporal modeling of visual, audio, and other modalities within a shared latent space. Methodologically, we introduce a unified multimodal latent-space alignment network coupled with hybrid relative and absolute temporal encoding, and develop a scalable data synthesis pipeline covering 24 million single- and multimodal dialogues. Trained on only 0.2 trillion tokens, OmniAlignNet surpasses Qwen2.5-Omni by +19.05, +1.7, and +3.9 points on the DailyOmni, MMAR, and Video-MME benchmarks, respectively. Extensive evaluations further demonstrate strong generalization in real-world applications—including robotics, medical AI, and smart manufacturing—validating its practical efficacy across diverse domains.
This work addresses the limitations of existing multimodal approaches that rely on modality-specific encoders operating at disparate frame rates, which weakens cross-modal interaction and impedes fine-grained modeling of visual dynamics. To overcome this, the authors propose a unified Transformer backbone that symmetrically embeds audio and visual signals at a shared 25 fps within a common latent space, emulating human-like holistic perception. The architecture introduces three key innovations: the Omni-Encoder Token Template, Omni-RoPE, and Temporal Window Shifting, which jointly enable effective modality disentanglement while maintaining computational efficiency. The model significantly outperforms baseline methods on sign language recognition and fine-grained sports action analysis, and remains competitive on established benchmarks such as audio-visual question answering (AVQA) and speaker localization.
Video understanding and generation are challenging to unify due to their divergent objectives: the former requires compact semantic representations, while the latter demands fine-grained detail preservation and temporal consistency. To address this, this work proposes Vega, the first framework to jointly model both tasks within a unified architecture for video. Vega aligns textual and visual representations through a shared semantic vocabulary, employs an autoregressive model to predict semantic tokens of keyframes, and leverages these tokens to guide a diffusion model in generating high-resolution, temporally coherent video frames. This hybrid approach achieves state-of-the-art performance on both VBench (for generation) and VideoMME (for understanding), demonstrating the effectiveness and versatility of a unified framework for video understanding and generation.
Current video models exhibit limitations in temporal understanding and heavily rely on large-scale datasets with language supervision, resulting in high training costs and constrained concept learning. This work proposes motion as a core modality, introducing point trajectories—structured motion cues—as an independent input for the first time. By employing a masked autoencoder to reconstruct occluded trajectories, the method enables self-supervised video representation learning without requiring language annotations or extensive appearance-based data. This approach substantially enhances temporal perception and data efficiency. The learned TIME embeddings achieve state-of-the-art performance in zero-shot settings, using four orders of magnitude less training data than existing methods.
This work addresses the challenge of simultaneously achieving cross-modal generalization and preserving modality-specific characteristics in multimodal representation learning. To this end, we propose CoDAAR, a novel framework that constructs the first competition-free unified discrete representation space. CoDAAR leverages Discrete Temporal Alignment (DTA) and Cascaded Semantic Alignment (CSA) mechanisms to establish cross-modal semantic consensus while retaining modality uniqueness. Trained via a self-supervised reconstruction objective, the method overcomes inherent limitations of both continuous and discrete representation approaches. Extensive experiments demonstrate that CoDAAR achieves state-of-the-art performance across diverse tasks—including event classification, temporal localization, video segmentation, and cross-dataset transfer—establishing a new discrete paradigm for multimodal representation learning.
This work addresses the absence of a systematic evaluation benchmark for existing omnimodal embedding models, which hinders accurate assessment of their semantic alignment and cross-modal retrieval capabilities. To this end, we introduce MMEB-V3—the first comprehensive benchmark encompassing text, images, videos, audio, and agent-based scenarios—and construct OmniSET, a dataset of semantically equivalent tuples designed to disentangle true semantic similarity from modality-specific artifacts. Leveraging this framework, we uncover three critical issues for the first time: query modality bias, target modality mismatch, and failure of instruction-guided retrieval. Our experiments reveal that state-of-the-art models exhibit significant deficiencies in adhering to modality constraints and achieving symmetric cross-modal retrieval, thereby providing a diagnostic foundation and clear directions for future research in omnimodal representation learning.