graph-aware ntk coupling

Designs and analyzes methods that couple neural tangent kernels with model representations that are explicitly aware of graph structure and language-model embeddings. Concretely, builds graph-aware NTK approximations and alignment procedures to approximate GNN (and LM–GNN hybrid) learning dynamics and enable kernel-based prediction or representation transfer that captures both structural and textual signals without repeatedly retraining LM–GNN models.

graph-awarentkcoupling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.9
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Simplifying Graph Neural Kernels: from Stacking Layers to Collapsed Structure

Jul 04, 2025
LW
Lin Wang
🏛️ The Hong Kong Polytechnic University

Existing Graph Neural Tangent Kernels (GNTKs) bridge kernel methods and graph neural networks (GNNs), but their layer-wise stacking architecture incurs substantial redundant computation, resulting in high time complexity and poor scalability. Method: We propose the Simplified Graph Neural Kernel (SGTK/SGNK) framework, which replaces multi-layer stacking with continuous $K$-hop neighborhood aggregation, integrates a collapsed architecture design, and employs Gaussian process modeling to analytically compute activation expectations—bypassing iterative layer-wise propagation. Under the infinite-width graph network assumption, SGTK enables efficient high-order neighborhood modeling. Contribution/Results: Theoretically and empirically, SGTK achieves comparable accuracy to GNTK on both node and graph classification tasks, while significantly reducing time complexity. It markedly improves computational efficiency and scalability, enabling practical deployment on larger graphs without sacrificing representational power.

Models infinitely wide GNNs as Gaussian Processes efficientlyReduces redundant computations in Graph Neural Tangent Kernel (GNTK)Simplifies kernel computation via continuous K-step aggregation

Existing Neural Tangent Kernel (NTK) theory relies on convergence bounds derived from the smallest eigenvalue, which are overly pessimistic and fail to explain the rapid convergence observed in practical neural network training. This work proposes a refined analytical framework based on the alignment between Label-NTK and Residual-NTK, revealing for the first time that the projections of labels and training residuals onto NTK eigenvectors scale proportionally with their corresponding eigenvalues. Leveraging this insight, the authors derive tight convergence and improved generalization bounds that depend on the full spectral structure of the NTK. Combining NTK linearized dynamics, spectral analysis, theoretical proofs, and extensive experiments across multiple datasets—including both MLPs and CNNs—the proposed bounds significantly outperform classical worst-case results, more accurately capture real-world training dynamics, and validate theoretical predictions on standard benchmarks.

convergence boundeigenvalue spectrumNeural Tangent Kernel

Depth-induced NTK: Bridging Over-parameterized Neural Networks and Deep Neural Kernels

Nov 05, 2025
YT
Yong-Ming Tian
🏛️ Nanjing University | City University of Hong Kong

Existing NTK theory primarily applies to infinite-width networks, neglecting the impact of depth on representation learning. This work systematically incorporates network depth as an explicit variable, proposing the Depth-driven Neural Tangent Kernel (D-NTK): finite-depth networks are mapped to Gaussian processes via skip connections, ensuring kernel convergence as depth tends to infinity. Theoretically, we characterize D-NTK’s dynamic stability and spectral properties, proving its training invariance and ability to suppress feature collapse. By integrating functional-space analysis with spectral decomposition, we establish a rigorous theoretical linkage among depth, kernel behavior, and generalization. Empirically, D-NTK consistently outperforms standard NTK on image classification and regression tasks, demonstrating enhanced expressive power and generalization performance—thereby bridging the theoretical gap between infinite-width and finite-depth regimes.

Analyzing training invariance and spectrum to stabilize kernel dynamicsBridging over-parameterized neural networks with deep neural kernels theoreticallyExtending NTK beyond infinite-width to incorporate network depth effects

This work proposes the Distilled Neural Tangent Kernel (DNTK), a novel approach that integrates dataset distillation into the input space of the Neural Tangent Kernel (NTK) to address its high computational cost stemming from large Jacobian matrices. By combining Jacobian projection with low-rank approximation, DNTK substantially reduces computational complexity while preserving the kernel structure and predictive performance. Theoretical analysis and empirical results demonstrate that NTK matrices across various architectures exhibit low effective rank, which can be effectively retained through distillation. The method achieves up to five orders of magnitude reduction in NTK computation overhead and decreases Jacobian evaluation costs by 20–100×, striking a favorable balance between efficiency and fidelity.

computational complexitydataset distillationJacobian computation

Graph Kernel Neural Networks

Dec 14, 2021
LC
Luca Cosmo
🏛️ Ca' Foscari University of Venice | Oxford University | Sapienza University of Rome | The Hong Kong Polytechnic University

Standard convolutional operations cannot be directly applied to graph-structured data due to its irregular, non-Euclidean topology. Method: This paper proposes Graph Kernel-driven Learnable Structural Convolution (GK-Conv), a purely structural, end-to-end modeling framework operating directly on non-Euclidean graph domains. GK-Conv eliminates explicit graph embedding and instead constructs a parameterized, structural convolutional operator grounded in generic graph kernel functions—enabling plug-and-play integration of arbitrary graph kernels and generating CNN-style, interpretable structural masks. The model is fully differentiable and optimized via ablation-guided hyperparameter analysis. Contribution/Results: GK-Conv achieves state-of-the-art performance across multiple graph classification and regression benchmarks, empirically validating the central claim that strong generalization can be attained using topology alone—without node or edge features.

Developing graph kernels for structural learning without embeddingsExtending convolution operators to irregular graph structuresProviding interpretable structural masks through graph kernel networks

Latest Papers

What's happening recently
View more

Existing graph neural tangent kernels model only pairwise relationships, limiting their ability to capture higher-order interactions and topological structures. This work introduces Hodge theory into neural tangent kernels for the first time, proposing an infinite-width kernel method tailored to edge features on simplicial complexes. By jointly modeling vertex-sharing and filled-simplex coupling through upper and lower Hodge interactions, the approach enables topology-aware learning of higher-order structures. Grounded in Hodge decomposition, the method reveals interpretable learning dynamics associated with gradient, harmonic, and curl components, integrating simplicial message passing, spectral analysis, and stability theory. Experiments on synthetic tasks and high-order link prediction on DBLP demonstrate substantial improvements in expressivity, learning efficiency, and predictive performance.

graph neural tangent kernelhigher-order interactionsHodge theory

This work addresses the intractability of computing the global empirical Neural Tangent Kernel (NTK) in finite-width neural networks, which hinders a precise understanding of gradient descent dynamics. Viewing model states as solutions to implicit constraints, the authors propose an operator factorization framework that decomposes the NTK into a product of a parameter–state interaction operator \( K \) and a state–state dependency operator \( P \). They establish a general Kronecker kernel theorem, proving that \( K \) admits an exact representation as a function of the Gram matrix of weight sites, thereby revealing the intrinsic low-rank structure of the NTK and its induced learning bias. Leveraging implicit modeling, matrix-free randomized linear algebra, and kpflow implementation, the method enables efficient NTK computation and demonstrates that architectures such as RNNs and Transformers inherently yield low-rank NTKs, clarifying how gradient descent favors dominant modes and how initialization constrains task-specific learning capacity.

finite-width networksgradient descentlow-rank

This work proposes RAMP, a novel approach that addresses the limitations of existing methods which compress node text in text-rich graphs into static embeddings, thereby causing information loss and decoupling structural reasoning from original semantics. RAMP uniquely integrates a large language model (LLM) as a graph-native message aggregation operator, dynamically anchoring raw textual content during message passing and refining neighbor messages to achieve deep integration of structural propagation and contextual text understanding. Departing from conventional feature extraction paradigms, RAMP introduces a dual-representation mechanism grounded in raw text, enabling unified support for both discriminative and generative tasks. Extensive experiments demonstrate that RAMP achieves state-of-the-art performance across multiple text-rich graph benchmarks, effectively bridging the gap between graph-structured message passing and deep textual reasoning.

graph learninginformation bottleneckLLM

This work addresses the challenge of extending Neural Tangent Kernel (NTK) theory—originally developed for regression—to classification settings, where cross-entropy loss typically drives logits to diverge, thereby violating the linearization assumption underpinning NTK. By introducing either parameter-space regularization or non-degenerate target conditions, the paper establishes, for the first time, sufficient conditions under which sufficiently wide neural networks maintain a "lazy training" regime in classification tasks, ensuring the NTK remains approximately constant throughout training. This advancement enables a rigorous extension of NTK theory to classification, allowing precise characterization of both training dynamics and generalization behavior. Moreover, it reveals a theoretical connection between the predictive distribution induced by random initialization and Bayesian inference.

classificationcross-entropy losslazy training regime

This work addresses a critical yet overlooked issue in graph neural networks (GNNs): during link prediction training, mini-batch sampling can introduce class composition bias, which—particularly when combined with batch normalization—leads models to learn spurious heuristics. Consequently, the learned representations become misaligned with those beneficial for node classification, challenging the common assumption that GNNs produce task-agnostic, transferable embeddings. The study systematically uncovers this bias for the first time and proposes a correction mechanism that aligns model representations with the intrinsic structural properties of the graph. Experimental results demonstrate that the corrected models focus more on class-relevant features, suggesting that standard training protocols may substantially overestimate the generalization capability of link predictors.

batch normalizationgraph neural networksgraph representation