Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation

📅 2025-03-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Real-world multi-view data often exhibit heterogeneity and incompleteness, undermining the robustness and generalizability of existing methods. To address this, we propose an unsupervised robust multi-view learning framework. First, we design a sample-level attention mechanism to adaptively fuse heterogeneous view representations. Second, we introduce simulated-perturbation contrastive learning within a dual-path architecture to align representations under noise perturbations, enabling dynamic noise modeling. Third, we establish a synergistic optimization paradigm integrating representation fusion and alignment. The framework is fully unsupervised and compatible with multi-view Transformers and cross-modal hashing retrieval. Extensive experiments demonstrate state-of-the-art performance on unsupervised clustering, noisy-label classification, and cross-modal hashing retrieval tasks. Ablation studies validate the efficacy of each component.

Technology Category

Machine Learning: Multi-instance/Multi-view LearningComputer Vision: Multi-modal VisionIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphs
📝 Abstract
Recently, multi-view learning (MVL) has garnered significant attention due to its ability to fuse discriminative information from multiple views. However, real-world multi-view datasets are often heterogeneous and imperfect, which usually makes MVL methods designed for specific combinations of views lack application potential and limits their effectiveness. To address this issue, we propose a novel robust MVL method (namely RML) with simultaneous representation fusion and alignment. Specifically, we introduce a simple yet effective multi-view transformer fusion network where we transform heterogeneous multi-view data into homogeneous word embeddings, and then integrate multiple views by the sample-level attention mechanism to obtain a fused representation. Furthermore, we propose a simulated perturbation based multi-view contrastive learning framework that dynamically generates the noise and unusable perturbations for simulating imperfect data conditions. The simulated noisy and unusable data obtain two distinct fused representations, and we utilize contrastive learning to align them for learning discriminative and robust representations. Our RML is self-supervised and can also be applied for downstream tasks as a regularization. In experiments, we employ it in unsupervised multi-view clustering, noise-label classification, and as a plug-and-play module for cross-modal hashing retrieval. Extensive comparison experiments and ablation studies validate the effectiveness of RML.
Problem

Research questions and friction points this paper is trying to address.

Addresses heterogeneity and imperfections in multi-view datasets.
Proposes robust multi-view learning via representation fusion and alignment.
Enhances discriminative and robust representations using simulated perturbations.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-view transformer fusion network for homogeneous embeddings
Sample-level attention mechanism for representation fusion
Simulated perturbation framework for robust contrastive learning
🔎 Similar Papers
2024-07-16arXiv.orgCitations: 3
J
Jie Xu
University of Electronic Science and Technology of China, Chengdu, China; Singapore University of Technology and Design, Singapore
N
Na Zhao
Singapore University of Technology and Design, Singapore
G
Gang Niu
Southeast University, Nanjing, China
Masashi Sugiyama
Masashi Sugiyama
Director, RIKEN Center for Advanced Intelligence Project / Professor, The University of Tokyo
Machine LearningData MiningArtificial Intelligence
X
Xiaofeng Zhu
University of Electronic Science and Technology of China, Chengdu, China; Hainan University, Haikou, China