Multi-Class, Multi-Tier Network Intrusion Detection: A Comprehensive and Reproducible Benchmark

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unreliability of existing intrusion detection system (IDS) benchmarks, where labeling errors and evaluation biases lead to inflated performance claims. To mitigate this, we propose a nested multi-level evaluation framework built upon a label-corrected version of the CIC-IDS2017 dataset. The framework establishes a reproducible evaluation pipeline spanning binary to fine-grained classification tasks, with extensible support for feature extraction and new model integration. Methodologically, we employ an ensemble learning approach combining Random Forest, XGBoost, and LightGBM via soft voting, complemented by feature importance analysis. Experimental results demonstrate that the proposed ensemble achieves macro-F1 scores of 0.955, 0.980, and 0.999 on fine-grained, coarse-grained, and binary classification tasks, respectively. This work thereby establishes a reliable, scalable, and open-source benchmark for future IDS research.
📝 Abstract
Machine learning (ML) and deep learning (DL) have dominated Intrusion Detection System (IDS) research in recent years. Unfortunately, many existing studies have produced inflated results and unreliable benchmarks due to critical oversights and mistakes in the ML and DL pipeline, from data collection and labeling to feature engineering and model training and evaluation. CIC-IDS2017 is a standard benchmark for network intrusion detection. Still, many published results on this dataset are difficult to compare due to labeling errors, inconsistent flow extraction, potential leakage, and performance evaluation metrics dominated by benign traffic. In this paper, we present a comprehensive benchmark with corrected PCAP-level labeling and a complete evaluation pipeline with diverse ML models. We evaluate eleven tabular classifiers at three nested levels: binary attack detection, nine-class attack-family attribution, and fifteen-class fine-grained classification. A soft-voting ensemble of Random Forest, XGBoost, and LightGBM obtains the best fine-tier macro-F1 of 0.955, with coarse and binary macro-F1 scores of 0.980 and 0.999, respectively. We further conducted a feature selection study based on an analysis of feature importance. This comprehensive benchmark pipeline is configurable and open-source, enabling new feature extraction and model plugins for new datasets. Future work should use this pipeline as a reference point for richer features, rare-class analysis, and model generalization towards new datasets and attack classes.
Problem

Research questions and friction points this paper is trying to address.

Network Intrusion Detection
Benchmark Reproducibility
Evaluation Pipeline
Data Labeling Errors
Machine Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Network Intrusion Detection
Benchmark Pipeline
Multi-tier Classification
Ensemble Learning
Feature Selection
Y
Yufeng Xin
RENCI, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA
B
Bryant Goseland
RENCI, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA
Mohamed Rahouti
Mohamed Rahouti
Fordham University
Computer networking and securityblockchain technologyAI and machine learning