🤖 AI Summary
This study addresses the limitations of cross-view interaction and information distortion in asynchronous multi-view fusion by proposing a synchronous neural diffusion paradigm. The proposed method models the multi-view feature space as a unified diffusion dynamical system, enabling the synchronous and adaptive integration of both intra-view and inter-view information. Furthermore, it introduces energy-topological sampling and an Ego-Net architecture to effectively balance model expressiveness with computational efficiency. Extensive experiments on real-world datasets demonstrate that this approach significantly outperforms existing baseline models while maintaining both theoretical elegance and computational efficiency.
📝 Abstract
Multi-view learning seeks to learn more comprehensive representations by exploiting the complementarity and consistency across diverse modalities or views. However, existing multi-view fusion strategies treat intra- and inter-view fusion as independent stages, without simultaneously considering the evolution within views and the dependency across views. Such an asynchronous fusion paradigm inevitably constrains cross-view interactions due to conflicting view-specific structural inductive biases. As a result, information flow is prone to distortion and compression along intermediate pathways, confining the model to learn within a restricted solution space. To address this, we propose Synchronous Multi-view Neural Diffusion (SynMDiff), which conceptualizes the multi-view feature space as a unified dynamical system driven by a diffusion process. By modeling the diffusion flow across arbitrary dyadic feature interactions in a joint space, SynMDiff enables the concurrent and adaptive intra- and inter-view information fusion. While a direct implementation of this synchronized mechanism incurs prohibitive computational costs, we further introduce an energy-based topological sampling strategy and an Ego-Net style centralized training architecture, ensuring both efficiency and scalability during learning and inference. Due to its conceptual elegance and computational efficacy, evaluations on real-world datasets demonstrate that SynMDiff outperforms the baselines by a large margin.