🤖 AI Summary
Existing interpretable intrusion detection systems (IDS) commonly rely on post-hoc approximators coupled with black-box classifiers, resulting in non-auditable rules, incomplete feature representation, and potential misinterpretation. To address this, we propose an end-to-end interpretable IDS framework: first, a multi-granularity Gaussian discretization method models continuous features in a human-readable, semantically meaningful manner; second, leveraging Interpretable Generalization, the framework directly learns auditable logical rules that distinguish benign from malicious traffic without approximation. The method achieves high accuracy and full transparency—even under extremely low training sample regimes—eliminating reliance on post-hoc surrogate models. Evaluated across nine distinct splits of the UKM-IDS20 dataset, it attains an average precision gain of ≥4 percentage points over state-of-the-art interpretable baselines, while maintaining near-perfect recall (≈1.0). These results demonstrate superior few-shot generalization capability and cross-dataset robustness.
📝 Abstract
Explainable intrusion detection systems (IDS) are now recognized as essential for mission-critical networks, yet most "XAI" pipelines still bolt an approximate explainer onto an opaque classifier, leaving analysts with partial and sometimes misleading insights. The Interpretable Generalization (IG) mechanism, published in IEEE Transactions on Information Forensics and Security, eliminates that bottleneck by learning coherent patterns - feature combinations unique to benign or malicious traffic - and turning them into fully auditable rules. IG already delivers outstanding precision, recall, and AUC on NSL-KDD, UNSW-NB15, and UKM-IDS20, even when trained on only 10% of the data. To raise precision further without sacrificing transparency, we introduce Multi-Granular Discretization (IG-MD), which represents every continuous feature at several Gaussian-based resolutions. On UKM-IDS20, IG-MD lifts precision by greater than or equal to 4 percentage points across all nine train-test splits while preserving recall approximately equal to 1.0, demonstrating that a single interpretation-ready model can scale across domains without bespoke tuning.