Score
Designs and implements algorithms and models that learn representations and predictors from multiple complementary views or feature sets (e.g., different modalities, sensors, or feature groups), including view-specific encoders, joint/fused representations, alignment and consistency objectives, and fusion strategies to improve predictive accuracy and robustness. Analyzes and evaluates methods for cross-view alignment, handling missing or noisy views, and how fusion or aggregation choices affect downstream tasks and generalization.
Multi-view learning lacks theoretical foundations for generalization, particularly in predicting model performance on unseen scenarios involving reconstruction and classification. Method: We establish the first information-theoretic unified generalization error bound, explicitly characterizing how consensus and complementary information influence representation disentanglement and generalization. We propose an information-bottleneck-regularized framework with provable generalization guarantees, designing data-dependent leave-one-out and supersample bounds, and deriving fast-rate convergence bounds in the interpolation regime. Contribution/Results: Our theoretical analysis yields a tight, computationally tractable upper bound on generalization error, revealing the synergistic interplay between consensus and complementary information. Empirical evaluation demonstrates strong correlation between the derived bound and actual generalization gaps, significantly enhancing interpretability and predictive capability for multi-view models.
Real-world multi-view data often exhibit heterogeneity and incompleteness, undermining the robustness and generalizability of existing methods. To address this, we propose an unsupervised robust multi-view learning framework. First, we design a sample-level attention mechanism to adaptively fuse heterogeneous view representations. Second, we introduce simulated-perturbation contrastive learning within a dual-path architecture to align representations under noise perturbations, enabling dynamic noise modeling. Third, we establish a synergistic optimization paradigm integrating representation fusion and alignment. The framework is fully unsupervised and compatible with multi-view Transformers and cross-modal hashing retrieval. Extensive experiments demonstrate state-of-the-art performance on unsupervised clustering, noisy-label classification, and cross-modal hashing retrieval tasks. Ablation studies validate the efficacy of each component.
This work addresses representation learning in distributed multi-view settings, where agents must extract features from local views alone—without explicit coordination—such that the union of these representations is both sufficient and necessary for decoding joint labels. We derive a novel data-dependent, symmetric prior-based generalization bound (in the Minimum Description Length sense), theoretically showing that Gaussian product mixture priors naturally encourage redundant feature extraction. Leveraging this insight, we propose a weighted attention mechanism. Our method integrates relative entropy-based generalization analysis, variational inference, and multi-view collaborative regularization. Experiments demonstrate: (i) superior performance over VIB and CDVIB on single-view tasks; (ii) significant generalization gains from controlled redundancy in multi-view settings; and (iii) strong alignment between theoretical guarantees and empirical results.
RGB-D multimodal fusion mechanisms have long suffered from poor interpretability, and the fundamental nature of cross-modal complementarity remains unclear. Method: This paper establishes the first interpretability analysis framework specifically targeting the fusion process, introducing a joint metric of semantic variance and feature similarity to systematically characterize cross-modal representation consistency, intra-modal evolutionary patterns, and collaborative optimization logic. Through cross-layer feature comparison and quantitative semantic analysis, we identify a prevalent imbalance between consistency and specificity in mainstream fusion strategies. Contribution/Results: We formalize a “specificity-driven inference under consistency constraints” principle that explicates cross-modal complementarity. Our framework provides both theoretical foundations and a verifiable evaluation paradigm for designing trustworthy, generalizable multimodal fusion models.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
This study addresses the absence of a unified multi-view learning toolchain within the Python ecosystem by proposing polyview, a scikit-learn-compatible library. Built upon a composable architecture centered around core classes, polyview provides unified interfaces for embedding, clustering, fusion, and missing-view handling. It facilitates the flexible composition of heterogeneous workflows and enables seamless transitions between multi-view and single-view processing stages. Experimental evaluations on five real-world datasets demonstrate that its components outperform existing baseline libraries. By filling the gap in end-to-end multi-view learning tools, this work offers an efficient and practical foundational framework for benchmarking and prototyping.
研究解决了多视图融合中性能随编码器数量非单调变化的问题,提出KAGES方法选择任务对齐的紧凑视图集,提高下游任务表现。
This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.
Cross-view image matching remains challenging due to significant viewpoint discrepancies, compounded by the absence of a unified problem formulation, model architecture, and evaluation protocol in the field. This work presents a systematic survey of the area and introduces the first structured taxonomy encompassing feature extraction, uni- and multimodal matchers, the integration of vision foundation models, and robust training strategies. Under a consistent experimental protocol, the study conducts a fair benchmark evaluation of prominent methods, revealing key design principles underlying the evolution from task-specific models toward generalizable correspondence frameworks. Furthermore, it establishes a reproducible evaluation platform and identifies critical future directions, including computational efficiency, robustness under extreme conditions, and cross-domain generalization.
This work addresses the challenge of arbitrary modality missingness during both training and inference in real-world multimodal scenarios—a setting where existing methods often fail due to their reliance on predefined missing patterns. The paper proposes a novel multimodal collaborative learning framework that abandons conventional fusion strategies and instead introduces, for the first time, a synergistic mechanism combining feature-level knowledge transfer with decision-level consistency constraints, explicitly designed to handle any missing modality configuration. Notably, the approach makes no assumptions about missingness patterns and adaptively processes any subset of input modalities. Experiments on two multimodal classification benchmarks demonstrate substantial robustness gains: the model not only outperforms competitors under single-modality missing conditions but also maintains strong performance even when only a single modality remains available.