fuse heterogeneous data

Designs, implements, or evaluates pipelines, networks, and methods that combine heterogeneous inputs — different modalities, feature sets, temporal frames, spatial scales, or separate models — into unified feature representations or fused outputs (including feature-level, model-level, multi-scale, multi-frame, and multi-source fusion). This work covers aligning and normalizing disparate formats and schemas, grouping and matching features, applying feature perturbation and weighting schemes, handling real and complex-valued features, and developing fusion strategies and techniques that preserve modality-specific signals while supporting downstream analysis or prediction.

fuseheterogeneousdata

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$194K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Enhancing Multimodal Unified Representations for Cross Modal Generalization

Mar 08, 2024
HH
Hai Huang
🏛️ Zhejiang University | Huawei

Existing approaches to enhancing the interpretability of multimodal unified representations rely on discretized representations but suffer from two key limitations: (1) Euclidean distance-based quantification ignores dimensional heterogeneity, inducing representation redundancy; and (2) uniform cross-modal alignment neglects modality-specific characteristics. To address these issues, we propose Training-Free Codebook Optimization (TOC) and Fine-Grained/Coarse-Grained Inter-Modal Information Decoupling (FCID)—the first framework enabling post-pretraining, gradient-free representation refinement and modality-adaptive information decoupling. TOC mitigates quantization redundancy via unsupervised codebook refinement, while FCID explicitly models modality-specific properties and disentangles shared versus private cross-modal information. Evaluated on cross-modal retrieval and zero-shot transfer tasks, our method achieves significant improvements over state-of-the-art baselines: representation redundancy is reduced by 37%, and modality specificity is enhanced by 21%.

Addressing Euclidean distance quantization limitations in feature dimensionsImproving multimodal representation interpretability with discrete unified methodsOptimizing cross-modal alignment by leveraging unique modality characteristics

Unsupervised Multi-modal Feature Alignment for Time Series Representation Learning

Dec 09, 2023
CL
Chen Liang
🏛️ Harbin Institute of Technology

To address the challenges of complex feature fusion and poor scalability in unsupervised multimodal time-series representation learning, this paper proposes a spectral-graph-theoretic framework for implicit cross-modal alignment. Methodologically, it abandons explicit multimodal feature fusion and instead adopts a single-encoder, multi-view architecture: diverse time-series views—such as frequency-domain, image-based, and symbolic representations—are generated via modality-agnostic temporal transformations, and spectral-graph-guided contrastive loss enables unsupervised cross-modal alignment. This lightweight design implicitly captures latent inter-modal dependencies, strengthening inductive bias and generalization capacity. Extensive experiments across multiple domains demonstrate state-of-the-art performance on downstream tasks—including classification, forecasting, and anomaly detection—surpassing existing unsupervised methods by an average accuracy gain of 3.2%–7.8%.

Aligning multi-modal time series features for representation learningImproving inductive bias in unsupervised time series encodingReducing feature fusion complexity to enhance scalability

Current biomedical multimodal models face critical bottlenecks: reliance on end-to-end training, exponential growth in computational complexity with modality count, severe performance degradation under extreme modality imbalance, and rigid topological coupling. To address these, we propose MM-Lego—a tuning-free, universal multimodal fusion framework. It introduces a novel frequency-domain feature harmonization mechanism that achieves shape alignment and interference-free merging of arbitrary unimodal encoders. We further design modality-agnostic wrappers and zero-/few-shot model merging strategies, enabling topology-agnostic fusion and robust modeling under highly imbalanced modalities. Crucially, MM-Lego requires no fine-tuning yet matches or surpasses end-to-end models in performance, while maintaining full encoder compatibility. Evaluated across seven biomedical benchmark datasets, it achieves state-of-the-art results on five—demonstrating unprecedented flexibility, efficiency, and generalizability in biomedical multimodal learning.

Enabling flexible model merging with minimal fine-tuningHandling diverse data modalities in biomedical machine learningOvercoming limitations of existing multimodal fusion approaches

LEGO: Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion

Oct 02, 2024
DD
Dexuan Ding
🏛️ Australian National University | Data61/CSIRO | Curtin University

To address weak structural modeling, shallow cross-modal interactions, difficult alignment, and poor interpretability in fusing heterogeneous multimodal features—spanning domains, granularities (e.g., token, patch, frame, clip), and modalities—this paper proposes a relation-centered, learnable graph-power fusion paradigm. It maps high-dimensional features into an interpretable graph space and constructs cross-granularity relational graphs. A learnable graph-power operator is introduced to aggregate element-wise relational scores via multivariate polynomials over homogeneous graphs, enabling structural-aware deep interaction. The method balances expressive power and interpretability, achieving multimodal fusion (text, image, video) without explicit alignment. Evaluated on video anomaly detection, it significantly outperforms concatenation, attention-based, and conventional nonlinear fusion baselines, demonstrating strong generalization and effectiveness.

Captures deep feature interactions in graph spaceImproves video anomaly detection across domainsLearnable graph fusion for multi-modal features

Interpretation on Multi-modal Visual Fusion

Aug 19, 2023
HC
Hao Chen
🏛️ Southeast University | Beijing University of Technology

RGB-D multimodal fusion mechanisms have long suffered from poor interpretability, and the fundamental nature of cross-modal complementarity remains unclear. Method: This paper establishes the first interpretability analysis framework specifically targeting the fusion process, introducing a joint metric of semantic variance and feature similarity to systematically characterize cross-modal representation consistency, intra-modal evolutionary patterns, and collaborative optimization logic. Through cross-layer feature comparison and quantitative semantic analysis, we identify a prevalent imbalance between consistency and specificity in mainstream fusion strategies. Contribution/Results: We formalize a “specificity-driven inference under consistency constraints” principle that explicates cross-modal complementarity. Our framework provides both theoretical foundations and a verifiable evaluation paradigm for designing trustworthy, generalizable multimodal fusion models.

Analyzing feature consistency and specialty across RGB-D modalitiesDeveloping improved fusion strategies for multi-modal RGB-D learningUnderstanding complementary and fusion mechanisms in RGB-D models

Latest Papers

What's happening recently
View more

This work addresses the challenge that certain modalities may introduce interference under specific inputs in multimodal fusion, thereby degrading model performance. To mitigate this issue, the authors propose a plug-and-play pre-fusion calibration module that leverages cross-modal summary contrast to extract supportive and conflicting cues, generating instance-level and dimension-level modulation signals. These signals dynamically enhance beneficial features while suppressing misleading information. The method achieves, for the first time, fine-grained, conflict-aware modulation prior to fusion and is compatible with both sequential and convolutional architectures. It consistently improves performance across five benchmark tasks—including emotion understanding and action recognition—and demonstrates enhanced robustness and consistency under modality missingness and data perturbations.

contextual adjustmentcross-modal conflictmodality calibration

This work addresses the limitations of existing approaches that predominantly rely on non-native late-fusion architectures and lack a systematic definition or unified framework for native multimodal modeling. The paper proposes a formal roadmap for Native Multimodal Modeling (NMM), introducing, for the first time, a rigorous formulation of “architectural nativeness.” Leveraging an input–output duality perspective, it categorizes models into three types: Multi-to-Text, Multi-to-Target, and Multi-to-Multi. A full-stack, industrial-grade NMM framework is developed, encompassing data governance, early/mid-stage fusion, end-to-end training, and deployment, all realized through a unified Transformer architecture that enables symbiotic cross-modal understanding and generation. This paradigm demonstrates superior performance and strong scalability across multiple tasks, offering a clear pathway toward truly native multimodal models.

architectural nativitymultimodal architecturemultimodal fusion

This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.

concatenationcross-attentionfeature alignment

This work addresses the challenge in multimodal sentiment analysis where modality-specific signal refinement and cross-modal interaction modeling often interfere with each other due to conflicting optimization objectives. To resolve this, the authors propose SeRIn, a novel architecture that decouples modality separation and cross-modal interaction into structured priors through a three-stage pipeline—separation, refinement, and integration—processing unimodal representations and cross-modal interactions via independent pathways before fusing them at the prediction stage. Notably, SeRIn adaptively adjusts modality weights without requiring explicit supervision. The method achieves state-of-the-art performance on the CH-SIMS and CMU-MOSEI benchmarks, yielding significant improvements across all evaluation metrics.

cross-modal interactionmodality-specific refinementmultimodal fusion

This work addresses the performance degradation in multimodal classification caused by modality imbalance by proposing deep ensembling as an alternative to explicit modality fusion, achieving effective multimodal classification through the combination of unimodal networks. The key contributions include the first demonstration that superior performance can be attained without explicit fusion, a heuristic strategy for allocating the number of ensemble models based on each modality’s predictive capability, and the construction of a controllable synthetic multimodal data framework with fitted scaling laws. Experiments show that, under identical parameter budgets, the proposed method significantly outperforms state-of-the-art late-fusion and intermediate-fusion approaches on both real-world and synthetic datasets, while the derived scaling laws reveal an asymptotic upper bound on ensemble performance.

deep ensembleslate-fusionmodality fusion

Hot Scholars

ZL

Zhen Lei

Associate Professor, OSCO Research Chair in Off-site Construction
Offsite ConstructionConstruction Engineering and Management
YG

Yulan Guo

Professor, Sun Yat-sen University
3D VisionMachine LearningRobotics
CG

Chenjuan Guo

Professor, East China Normal University
Data AnalyticsMachine Learning
HF

Huazhu Fu

Principal Scientist, IHPC, A*STAR
Medical Image AnalysisAI for HealthcareMedical AITrustworthy AI
XZ

Xiaoli Zhang

Jilin University
image fusiondata mining,image segmentation,deep learning