🤖 AI Summary
This study addresses the memory bottleneck in topological deep learning on large-scale graphs, where global domain construction renders training infeasible. To overcome this limitation, we propose Cluster-TNN, a framework that integrates graph pre-partitioning with runtime local dynamic sampling for mini-batch generation. Furthermore, it introduces the first local topology enhancement strategy that circumvents global materialization. Extensive evaluations demonstrate that our approach reduces peak GPU memory consumption by an average of 83.2% while maintaining competitive model performance. Notably, this work achieves, for the first time, scalable training of higher-order networks on large-scale datasets such as Reddit, thereby enabling practical topological deep learning applications on graphs previously considered computationally prohibitive.
📝 Abstract
Topological Deep Learning extends graph-based learning to higher-order domains, such as hypergraphs, cellular, and simplicial complexes. These domains are typically constructed from patterns in an input graph through a process of graph lifting. Full-domain training constructs and stores the complete lifted representation before model execution. On large and dense datasets like Reddit (233k nodes and 57.3M edges), this global materialization becomes a severe computational bottleneck, often rendering training infeasible. To address this limitation, we introduce Cluster-TNN, a domain-agnostic framework that avoids this bottleneck by lifting locally instead. After partitioning the input graph during preprocessing, at runtime Cluster-TNN dynamically samples groups of node clusters, reconstructs their induced subgraphs to form mini-batches, and applies the chosen lifting within each mini-batch. Retaining all edges among sampled nodes preserves the connectivity needed to construct higher-order structures across clusters, producing topological mini-batches that existing Topological Neural Networks can process directly. Across 21 matched comparisons with full-graph execution, Cluster-TNN reduces peak GPU memory in every configuration, by 83.2% on average while maintaining competitive predictive performance. Notably, such a reduction enables, to our knowledge, the first training of multiple different higher-order Topological Neural Networks on large datasets such as Reddit and OGBN Products. These results establish Cluster-TNN as a general strategy for scaling Topological Deep Learning beyond the limitations of global domain construction.