Score
Design and evaluate algorithms and training objectives that produce unified, view-invariant graph representations by aligning, fusing, or contrasting node- and edge-level features across multiple views or viewpoints. This includes methods to perform cross-view feature learning and alignment, disentangle entangled factors so embeddings preserve correspondence under viewpoint change, and to make representations generalize across different data sources or simulated vs. real distributions.
Link prediction on incomplete graphs suffers from difficulty in modeling graph representation invariance under structural perturbations. Method: This paper proposes a coupled graph augmentation and cross-view consistency learning framework. It generates two structurally complementary graph views via tailored augmentation strategies, jointly optimizes graph autoencoder reconstruction and cross-view contrastive loss, and incorporates a consistency regularization constraint to enforce robustness against structural perturbations. Contribution/Results: Theoretical analysis establishes convergence guarantees and formal invariance properties of the learned representations. Extensive experiments on multiple benchmark datasets demonstrate significant improvements over state-of-the-art methods in link prediction performance, validating the strong generalization capability and practical effectiveness of the proposed approach.
Existing graph classification methods exhibit limited out-of-distribution (OOD) generalization, primarily focusing on semantic invariance while neglecting the causal stability inherent in graph structure. Method: We propose a Unified Invariant Learning (UIL) framework—the first to jointly model structural invariance (enforced via graph-on-distance constraints on subgraph feature stability) and semantic invariance (achieved through environment partitioning and cross-environment contrastive learning for robust representation). UIL incorporates a theory-driven stable feature discrimination criterion and cross-environment consistency regularization to provably converge to causally stable graph features. Contribution/Results: UIL achieves significant improvements over state-of-the-art methods across multiple OOD graph classification benchmarks. We provide theoretical analysis proving its convergence advantage under causal stability assumptions. The implementation is publicly available.
To address weak structural modeling, shallow cross-modal interactions, difficult alignment, and poor interpretability in fusing heterogeneous multimodal features—spanning domains, granularities (e.g., token, patch, frame, clip), and modalities—this paper proposes a relation-centered, learnable graph-power fusion paradigm. It maps high-dimensional features into an interpretable graph space and constructs cross-granularity relational graphs. A learnable graph-power operator is introduced to aggregate element-wise relational scores via multivariate polynomials over homogeneous graphs, enabling structural-aware deep interaction. The method balances expressive power and interpretability, achieving multimodal fusion (text, image, video) without explicit alignment. Evaluated on video anomaly detection, it significantly outperforms concatenation, attention-based, and conventional nonlinear fusion baselines, demonstrating strong generalization and effectiveness.
This work addresses the limited expressive power of unsupervised graph alignment models, particularly in discriminating matching versus non-matching node pairs and enforcing structural matching constraints (e.g., bijectivity and mutual alignment). We propose CombAlign, a theoretically grounded hybrid framework that— for the first time—formalizes model expressivity from both discriminative capability and matching constraint perspectives. CombAlign integrates Gromov–Wasserstein optimal transport with Weisfeiler–Lehman-style node embedding, incorporates non-uniform marginal priors to encode structural biases, and refines alignments via maximum-weight bipartite matching. Evaluated on standard benchmarks, CombAlign achieves a 14.5% absolute improvement in alignment accuracy over state-of-the-art methods. Empirical results consistently validate the theoretical analysis, demonstrating strong alignment between expressivity characterization and practical performance.
Unlike vision and language domains, graph learning lacks a shared input space, as input features differ across graph datasets not only in semantics, but also in value ranges and dimensionality. This misalignment prevents graph models from generalizing across datasets, limiting their use as foundation models. In this work, we propose ALL-IN, a simple and theoretically grounded method that enables transferability across datasets with different input features. Our approach projects node features into a shared random space and constructs representations via covariance-based statistics, thus eliminating dependence on the original feature space. We show that the computed node-covariance operators and the resulting node representations are invariant in distribution to permutations of the input features. We further demonstrate that the expected operator exhibits invariance to general orthogonal transformations of the input features. Empirically, ALL-IN achieves strong performance across diverse node- and graph-level tasks on unseen datasets with new input features, without requiring architecture changes or retraining. These results point to a promising direction for input-agnostic, transferable graph models.
This work addresses the limitations of existing representation alignment methods, which predominantly rely on geometric properties and struggle to capture the global structural organization of model representations. To overcome this, the study introduces topological data analysis into the field for the first time, proposing a Mapper-based visual analytics framework. By integrating force-directed layout, Bubble Sets, motif querying, and membrane-inspired heuristics, the framework enables a unified analytical pipeline spanning global structure alignment, local region matching, and fine-grained pattern exploration. Case studies on language and multimodal models, complemented by expert evaluations, demonstrate that the approach effectively reveals and compares the topological organization of representations across different models or layers, offering deep structural insights.
This work addresses the limitations of existing unsupervised graph representation learning methods, which rely on the homophily assumption and thus struggle to model the heterogeneity of structure and attributes in heterophilic graphs, leading to loss of high-frequency information. To overcome this, the authors propose AlignGAE, a novel framework that employs dual encoders to separately capture structural and attribute information, incorporates node positional encoding to approximate neighborhood identity distributions, and introduces a dual reconstruction task—on both edges and attributes—to align complementary views. Notably, AlignGAE is the first to integrate a theoretically grounded neighborhood identity alignment strategy into spectral-aware learning, preserving view diversity while ensuring semantic consistency. Extensive experiments on twelve benchmark datasets demonstrate that AlignGAE achieves up to an 18.7% performance gain in node classification on heterophilic graphs while maintaining state-of-the-art results on homophilic graphs.
This work addresses the challenge in graph foundation models where heterogeneous node features across domains are difficult to unify, and naive dimension alignment often discards critical semantic information, limiting transferability. To overcome this, the authors propose SliGFM, a novel framework that establishes four guiding principles for graph feature unification: formal consistency, cross-domain transferability, information preservation, and backbone compatibility. Guided by these principles, SliGFM introduces a topology-aware sliding-window Transformer architecture that transforms heterogeneous features into ordered, fixed-dimensional semantic tokens through topological smoothness ranking, a shared sliding-window encoder, smoothness-aware attention mechanisms, and a generative reconstruction objective. Experiments demonstrate that SliGFM effectively captures transferable relational patterns while preserving original feature semantics, significantly enhancing the generalization performance of graph foundation models across diverse downstream tasks.