multimodal feature fusion

Designs and implements algorithms and model components that combine feature streams from multiple modalities—particularly audio/speech, text, and images—into unified representations for downstream prediction or generation. This includes building projection layers (e.g., audio-to-LLM), simple concatenation of embeddings, bidirectional cross-attention fusion, lightweight GRU-based sequence summaries, and sparse mixture-of-experts (top-k MoE) routing to control computation while preserving cross-modal cues.

multimodalfeaturefusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

Sep 26, 2025
CT
Changli Tang
🏛️ Tsinghua University | WeChat Vision | Tencent Inc.

To address the limited capability of multimodal large language models (MLLMs) in representing dynamic modalities such as audio and video, this paper introduces UniAV—the first MLLM-based universal audio-visual embedding framework—establishing a unified text-audio-video tri-modal embedding space. Methodologically, UniAV employs hierarchical cross-modal feature fusion, prompt-aware embedding generation, and multi-task joint training to achieve fine-grained semantic alignment and bidirectional retrieval and generation across arbitrary modalities. Its core contributions are: (1) the first LLM-driven, prompt-controllable, modality-agnostic unified embedding representation; and (2) state-of-the-art performance on the MMEB-v2 video understanding benchmark, along with significant improvements over prior methods in audio-video cross-modal retrieval and multimodal question answering.

Creating unified embeddings for text, audio, and videoEnabling any-to-any cross-modal retrieval between modalitiesGenerating prompt-aware embeddings for user instructions

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Jun 16, 2025
SZ
Shaolei Zhang
🏛️ Chinese Academy of Sciences | University of Chinese Academy of Sciences

Current large multimodal models (LMMs) rely on sequential concatenation and massive cross-modal data for alignment, resulting in inefficient interaction and limited flexibility. To address this, we propose a semantics-driven differential alignment strategy: vision–text alignment is performed via sequence-level concatenation, while speech–text alignment employs a CTC-guided layer-wise mapping for lightweight cross-modal transfer. We introduce the first layer-dimensional speech–text alignment mechanism, decoupling modality fusion strategies; and enable, for the first time, dual-path real-time intermediate outputs—ASR transcription and model response—during speech interaction. Built upon a large language model backbone, our architecture supports synchronous text–vision–speech interaction via multimodal joint training and inference. Experiments demonstrate state-of-the-art performance on visual understanding, speech interaction, and audio-visual joint tasks, significantly reducing speech-data dependency while enabling low-latency, high-fidelity real-time multimodal interaction.

Efficient alignment of text, vision, and speech modalitiesReducing data dependency for modality alignment learningSimultaneous interaction under various modality combinations

Understanding and Harnessing Sparsity in Unified Multimodal Models

Dec 02, 2025
SH
Shwai He
🏛️ ByteDance Seed | University of Maryland, College Park

Unified multimodal models suffer from inefficient inference due to mandatory full-component activation, yet systematic characterization of sparsity patterns across modules remains lacking. Method: We conduct the first systematic analysis of sparsity disparities between understanding and generation modules, revealing that the former is highly compressible while the latter is sensitive to compression. Building on this insight, we propose a dynamic-activation-pattern-based Mixture-of-Experts (MoE) adaptation framework: (i) training-free pruning for sparsity probing, integrated with joint depth-width compression analysis; and (ii) a sparse activation mechanism enabling expert freezing for fine-tuning and full-parameter training. Contribution/Results: Evaluated on the BAGEL model, our method achieves full-model performance while activating only ~50% of parameters—yielding substantial inference speedup without compromising generation quality, thus balancing efficiency and fidelity.

Current approaches lack systematic understanding of inefficiency distribution across model componentsGeneration components are highly sensitive to compression, causing sharp performance deteriorationUnified multimodal models suffer from inference inefficiencies due to unnecessary full model usage

This work addresses the limitation of existing multimodal large language models that employ shared parameters for speech and text, thereby neglecting their inherent representational differences and impairing modality-specific learning. To overcome this, the authors propose MoST, a novel modality-aware mixture-of-experts (MAMoE) architecture that assigns dedicated expert pathways for speech and text while incorporating shared experts to facilitate cross-modal fusion. As the first fully open-source speech-text mixture-of-experts large language model, MoST integrates modality-aware routing, post-training on open-source data, speech-text instruction tuning, and efficient alignment techniques. It achieves significant performance gains over comparable-scale models across diverse tasks—including automatic speech recognition, text-to-speech synthesis, audio language modeling, and spoken question answering. Ablation studies further confirm the effectiveness of both modality-specific routing and the shared expert mechanism.

cross-modal understandingMixture of Expertsmodality-specific representation

MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders

Sep 10, 2024
WZ
Wenyu Zhang
🏛️ Agency for Science, Technology and Research (A*STAR)

Existing AudioLLMs rely on a single pre-trained audio encoder, whose fixed representation capacity limits generalization across diverse audio tasks. To address this, we propose the Mixture of Weak Encoders (MoWE), which replaces the monolithic encoder with multiple lightweight audio encoders and employs a learnable gating mechanism to dynamically activate a task-adaptive subset of them—while keeping the large language model frozen. This enables low-overhead, task-specific feature enhancement without architectural or parametric overhead on the LLM. MoWE introduces, for the first time, a “mixture of weak encoders” paradigm that breaks the representational bottleneck by jointly optimizing expressiveness and efficiency: it incurs less than 3% additional parameters while significantly increasing representational diversity. Evaluated on a cross-domain multi-task audio understanding benchmark, MoWE achieves an average accuracy gain of +4.2% over strong single-encoder baselines, demonstrating superior adaptability and robustness.

Broaden applicability to diverse audio tasks.Enhance feature extraction in AudioLLMs.Improve multi-task performance with MoWE.

Latest Papers

What's happening recently
View more

This work addresses the lack of native support for multimodal generation—particularly multi-layer audio tokens in speech-language models—in existing high-throughput inference engines. Building upon vLLM, the authors propose a unified end-to-end inference pipeline for joint audio understanding and generation. They innovatively extend autoregressive decoding to enable delayed-mode deinterleaving and multi-stream collaborative sampling, while integrating a GPU-resident acoustic decoder for efficient waveform synthesis. By co-scheduling conditional and unconditional requests within continuous batching, the system efficiently supports Classifier-Free Guidance (CFG) with only a 20% throughput penalty relative to non-CFG execution, achieving 80% of the latter’s throughput efficiency. This approach significantly enhances both multimodal audio generation quality and overall system performance, and the complete framework is released as open source.

classifier-free guidancehigh-throughput inferencemulti-token prediction

This study addresses the limitations of large audio-language models in cross-domain representation modeling and multi-task understanding by proposing a unified audio encoder based on a Mixture-of-Experts (MoE) architecture. Methodologically, it introduces a SwiGLU shared expert decoupling network to integrate the strengths of mainstream encoders, alongside a novel two-stage instruction tuning strategy and Task-Specific Data Scaling (TSDS) technique to enhance model generalization. Experimental results demonstrate that the proposed model achieves state-of-the-art performance of 0.802 on the XARES-LLM benchmark and exhibits superior cross-domain understanding capabilities in the Interspeech 2026 Challenge.

Audio UnderstandingCross-domain Audio RepresentationLarge Audio Language Models

This work proposes ERNIE 5.0, the first trillion-parameter native autoregressive multimodal foundation model capable of unified processing of text, images, video, and audio. To address the challenge of efficient deployment under resource constraints, the model employs an ultra-sparse mixture-of-experts (MoE) architecture with a modality-agnostic expert routing mechanism and is trained from scratch using a unified “next group of tokens” prediction objective. A novel elastic training paradigm is introduced, enabling the simultaneous learning of a family of prunable submodels within a single pretraining run, with dynamic adjustment of depth, expert capacity, and sparsity. This approach systematically resolves the stability and efficiency challenges of multimodal reinforcement learning under ultra-sparse MoE settings, achieving balanced and state-of-the-art performance across both multimodal understanding and generation tasks.

autoregressive modelingelastic trainingmixture-of-experts

This study addresses the unclear mechanisms underlying cross-layer fusion of visual and textual information in multimodal large language models. We propose an architecture-aware diagnostic framework that systematically compares concatenation-based and native multimodal architectures. Through alignment decoupling, attention entropy analysis, intrinsic dimensionality estimation, causal intervention, and vision-specific Centered Kernel Alignment (CKA), we reveal how feature spaces are reorganized under different architectural paradigms. Our findings indicate that concatenation-based models exhibit a text-dominant fusion trajectory, whereas native models achieve early-stage vision-language co-adaptation. This work elucidates the distinct multimodal fusion mechanisms inherent to these two architectural paradigms, providing a principled theoretical foundation for future model design.

Architecture ParadigmsMechanistic InterpretabilityMultimodal Fusion

This work addresses the limitations of existing approaches that predominantly rely on non-native late-fusion architectures and lack a systematic definition or unified framework for native multimodal modeling. The paper proposes a formal roadmap for Native Multimodal Modeling (NMM), introducing, for the first time, a rigorous formulation of “architectural nativeness.” Leveraging an input–output duality perspective, it categorizes models into three types: Multi-to-Text, Multi-to-Target, and Multi-to-Multi. A full-stack, industrial-grade NMM framework is developed, encompassing data governance, early/mid-stage fusion, end-to-end training, and deployment, all realized through a unified Transformer architecture that enables symbiotic cross-modal understanding and generation. This paradigm demonstrates superior performance and strong scalability across multiple tasks, offering a clear pathway toward truly native multimodal models.

architectural nativitymultimodal architecturemultimodal fusion

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
YH

Yao Hu

浙江大学
Machine Learning
YR

Yiming Ren

Tsinghua University
Object Detection、Multimodal Large Language Model
MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
UB

Ulas Bagci

Northwestern University
artificial intelligencedeep learningbiomedical image analysismedical image computing