🤖 AI Summary
Existing Interpretability-Guided (IG) mechanisms for intrusion detection offer high interpretability but suffer from cubic time complexity and prohibitively large intermediate bitset volumes, rendering them infeasible to scale to full datasets on CPU. This work presents the first end-to-end, sampling-free GPU acceleration of the entire IG pipeline—including dense intersection and subset operations—implemented natively in PyTorch. Evaluated on the full NSL-KDD dataset, our approach achieves training in 18 minutes (116× speedup over CPU) while attaining Recall = 0.957, Precision = 0.973, and AUC = 0.961; single-flow inference completes in milliseconds. Our core contribution is the first GPU-native IG architecture explicitly designed for interpretable pattern discovery—uniquely reconciling formal interpretability guarantees with industrial-scale scalability.
📝 Abstract
The Interpretable Generalization (IG) mechanism recently published in IEEE Transactions on Information Forensics and Security delivers state-of-the-art, evidence-based intrusion detection by discovering coherent normal and attack patterns through exhaustive intersect-and-subset operations-yet its cubic-time complexity and large intermediate bitsets render full-scale datasets impractical on CPUs. We present IG-GPU, a PyTorch re-architecture that offloads all pairwise intersections and subset evaluations to commodity GPUs. Implemented on a single NVIDIA RTX 4070 Ti, in the 15k-record NSL-KDD dataset, IG-GPU shows a 116-fold speed-up over the multi-core CPU implementation of IG. In the full size of NSL-KDD (148k-record), given small training data (e.g., 10%-90% train-test split), IG-GPU runs in 18 minutes with Recall 0.957, Precision 0.973, and AUC 0.961, whereas IG required down-sampling to 15k-records to avoid memory exhaustion and obtained Recall 0.935, Precision 0.942, and AUC 0.940. The results confirm that IG-GPU is robust across scales and could provide millisecond-level per-flow inference once patterns are learned. IG-GPU thus bridges the gap between rigorous interpretability and real-time cyber-defense, offering a portable foundation for future work on hardware-aware scheduling, multi-GPU sharding, and dataset-specific sparsity optimizations.