🤖 AI Summary
This paper addresses the multi-class detection problem of five network intrusion types—Normal, DoS, Probe, R2L, and U2R—in the NSL-KDD dataset. We propose an efficient tree-ensemble-based detection framework, incorporating systematic feature analysis and standardized preprocessing (including categorical encoding and min-max normalization). Four supervised learning models—logistic regression, decision trees, random forest, and XGBoost—are comparatively evaluated. Experimental results demonstrate that random forest and XGBoost achieve superior performance, attaining ≈99% overall accuracy, high F1-scores, and strong AUC values—establishing a new performance benchmark for traditional network intrusion detection systems (NIDS). Furthermore, the study identifies persistent recognition bottlenecks for rare attack classes (R2L and U2R), revealing inherent challenges in long-tailed class imbalance. We propose targeted optimization strategies for minority-class detection, thereby contributing both methodological rigor and practical guidance for real-world NIDS deployment.
📝 Abstract
Intrusion detection is a critical component of cybersecurity, responsible for identifying unauthorized access or anomalous behavior in computer networks. This paper presents a comprehensive study on intrusion detection in networks using classical machine learning algorithms applied to the multiclass version of the NSL-KDD dataset (Normal, DoS, Probe, R2L, and U2R classes). The characteristics of NSL-KDD are described in detail, including its variants and class distribution, and the data preprocessing process (cleaning, coding, and normalization) is documented. Four supervised classification models were implemented: Logistic Regression, Decision Tree, Random Forest, and XGBoost, whose performance is evaluated using standard metrics (accuracy, recall, F1 score, confusion matrix, and area under the ROC curve). Experiments show that models based on tree sets (Random Forest and XGBoost) achieve the best performance, with accuracies approaching 99%, significantly outperforming logistic regression and individual decision trees. The ability of each model to detect each attack category is also analyzed, highlighting the challenges in identifying rare attacks (R2L and U2R). Finally, the implications of the results are discussed, comparing them with the state of the art, and potential avenues for future research are proposed, such as the application of class balancing techniques and deep learning models to improve intrusion detection.