Score
Designs and implements methods that generate two complementary augmented versions of a graph—by perturbing nodes, edges, or features or by altering substructures—to produce paired inputs for siamese/contrastive learning and to mitigate sparsity and noise while preserving essential relational semantics. Builds augmentation pipelines and evaluates how these graph-view transformations affect representation robustness and downstream task performance.
To address the distortion of graph properties and insufficient structural diversity caused by conventional data augmentation in graph classification, this work investigates graph spectral characteristics and, for the first time, identifies the critical role of low-frequency spectral components in conserving global graph properties—such as connectivity and clustering coefficient. We propose Dual-Prism (DP-Noise/DP-Mask), a spectral-aware augmentation framework grounded in graph Laplacian spectral decomposition. It imposes explicit constraints in the low-frequency subspace and synergistically integrates differentiable noise injection with masked reconstruction to jointly optimize property preservation and structural diversification. Evaluated on multiple standard graph classification benchmarks, our method improves GNN generalization performance by 2.3–5.1% on average and reduces bias in key topological properties by over 67%. The framework is both interpretable—owing to its spectral foundation—and practically deployable.
Graph neural networks (GNNs) for graph classification suffer from overfitting, while existing graph augmentation methods lack sufficient representation robustness. Method: This paper proposes an augmentation-aware representation learning framework. Its core innovations include: (i) modeling augmentation discrepancy as a learnable graph distance prediction task, jointly constraining structural and feature-level differences between augmented and original graphs; (ii) introducing an augmentation-aware supervised training paradigm that explicitly accounts for varying augmentation strengths—overcoming a key limitation of conventional contrastive learning; and (iii) incorporating multi-level consistency regularization to enhance classifier robustness against diverse augmented representations. Contribution/Results: The framework achieves state-of-the-art performance across supervised, semi-supervised, and cross-domain graph classification benchmarks, demonstrating significant improvements in generalization and robustness.
Traditional self-supervised learning overly relies on data augmentation while neglecting semantic relationships among instances. To address this, we propose the first systematic incorporation of graph-structured modeling to explicitly capture inter-instance associations. Specifically, we construct a teacher–student dual-stream k-nearest neighbor (k-NN) graph and integrate graph neural networks (GNNs) to enable multi-hop message passing, thereby unifying local augmentations with global contextual information. Our approach introduces a k-NN-based dual-stream architecture coupled with a representation refinement mechanism, breaking away from the conventional paradigm that learns solely from intra-instance variations. Extensive experiments demonstrate consistent improvements in linear evaluation accuracy: +7.3% on CIFAR-10, +3.2% on ImageNet-100, and +1.0% on ImageNet-1K—outperforming state-of-the-art methods. These results validate the effectiveness and generalizability of explicitly modeling inter-instance relationships for self-supervised representation learning.
Existing self-supervised graph representation learning methods predominantly rely on data perturbation for augmentation, neglecting semantic consistency and thereby limiting representation quality. To address this, we propose the Explanation-Preserving Augmentation (EPA) framework—the first to integrate graph explanation techniques (e.g., GNNExplainer variants) into graph augmentation design. EPA trains a lightweight explainer on a small set of labeled nodes to identify semantically critical substructures, then generates semantically consistent augmented graphs accordingly. Coupled with subgraph sampling, feature masking, and contrastive learning, EPA enables semantic-aware representation learning within semi-supervised GNNs. Theoretical analysis establishes its robustness to structural noise, while extensive experiments demonstrate that EPA consistently outperforms state-of-the-art semantic-agnostic methods across multiple benchmarks, significantly improving both semantic fidelity and generalization performance.
To address weak structural modeling, shallow cross-modal interactions, difficult alignment, and poor interpretability in fusing heterogeneous multimodal features—spanning domains, granularities (e.g., token, patch, frame, clip), and modalities—this paper proposes a relation-centered, learnable graph-power fusion paradigm. It maps high-dimensional features into an interpretable graph space and constructs cross-granularity relational graphs. A learnable graph-power operator is introduced to aggregate element-wise relational scores via multivariate polynomials over homogeneous graphs, enabling structural-aware deep interaction. The method balances expressive power and interpretability, achieving multimodal fusion (text, image, video) without explicit alignment. Evaluated on video anomaly detection, it significantly outperforms concatenation, attention-based, and conventional nonlinear fusion baselines, demonstrating strong generalization and effectiveness.
Unlike vision and language domains, graph learning lacks a shared input space, as input features differ across graph datasets not only in semantics, but also in value ranges and dimensionality. This misalignment prevents graph models from generalizing across datasets, limiting their use as foundation models. In this work, we propose ALL-IN, a simple and theoretically grounded method that enables transferability across datasets with different input features. Our approach projects node features into a shared random space and constructs representations via covariance-based statistics, thus eliminating dependence on the original feature space. We show that the computed node-covariance operators and the resulting node representations are invariant in distribution to permutations of the input features. We further demonstrate that the expected operator exhibits invariance to general orthogonal transformations of the input features. Empirically, ALL-IN achieves strong performance across diverse node- and graph-level tasks on unseen datasets with new input features, without requiring architecture changes or retraining. These results point to a promising direction for input-agnostic, transferable graph models.
This work addresses the challenge that graph structures naively derived from relational databases often suffer from information overload and semantic fragmentation, rendering them ill-suited for relational reasoning with graph neural networks (GNNs). To overcome this limitation, the authors propose an end-to-end structure optimizer that automatically constructs GNN-friendly relational graphs by jointly performing information filtering and semantic enrichment. The method uncovers key mechanisms underlying effective adaptation of relational graphs to GNN architectures. Empirical evaluation across 26 diverse tasks—including classification, regression, and recommendation—demonstrates consistent improvements in model accuracy, frequently accompanied by reduced inference overhead.