A Review of Various Datasets for Machine Learning Algorithm-Based Intrusion Detection System: Advances and Challenges

📅 2025-06-03
🏛️ Social Science Research Network
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses critical challenges in machine learning–based intrusion detection systems (IDS): dataset bias, misaligned evaluation metrics, and poor model generalizability. We conduct a systematic empirical analysis by performing the first large-scale, cross-classifier benchmark—evaluating ten mainstream classifiers (e.g., SVM, RF, XGBoost, ANN) across five widely used IDS datasets (KDDCUP’99, NSL-KDD, UNSW-NB15, CIC-IDS2017, CSE-CIC-IDS2018). Leveraging bibliometric analysis and tabular meta-analysis—including attack-type coverage, F1-score, and accuracy—we identify structural deficiencies in these datasets concerning attack representativeness, feature discriminability, and evaluation consistency. Based on these findings, we propose a dynamic dataset construction paradigm grounded in real-world network traffic characteristics. This paradigm supports the development of lightweight, highly generalizable next-generation IDS models. Our work delivers a reproducible methodological framework and practical guidelines for robust, evidence-based IDS research.

Technology Category

Machine Learning: Evaluation and AnalysisData Mining & Knowledge Management: Anomaly/Outlier DetectionMultiagent Systems: Adversarial Agents

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSecurity and Privacy: Large-scale security measurementsSocial Networks and Social Media: Fairness and bias in social network and social media analysis
📝 Abstract
IDS aims to protect computer networks from security threats by detecting, notifying, and taking appropriate action to prevent illegal access and protect confidential information. As the globe becomes increasingly dependent on technology and automated processes, ensuring secured systems, applications, and networks has become one of the most significant problems of this era. The global web and digital technology have significantly accelerated the evolution of the modern world, necessitating the use of telecommunications and data transfer platforms. Researchers are enhancing the effectiveness of IDS by incorporating popular datasets into machine learning algorithms. IDS, equipped with machine learning classifiers, enhances security attack detection accuracy by identifying normal or abnormal network traffic. This paper explores the methods of capturing and reviewing intrusion detection systems (IDS) and evaluates the challenges existing datasets face. A deluge of research on machine learning (ML) and deep learning (DL) architecture-based intrusion detection techniques has been conducted in the past ten years on various cybersecurity datasets, including KDDCUP'99, NSL-KDD, UNSW-NB15, CICIDS-2017, and CSE-CIC-IDS2018. We conducted a literature review and presented an in-depth analysis of various intrusion detection methods that use SVM, KNN, DT, LR, NB, RF, XGBOOST, Adaboost, and ANN. We provide an overview of each technique, explaining the role of the classifiers and algorithms used. A detailed tabular analysis highlights the datasets used, classifiers employed, attacks detected, evaluation metrics, and conclusions drawn. This article offers a thorough review for future IDS research.
Problem

Research questions and friction points this paper is trying to address.

Reviewing datasets for ML-based intrusion detection systems
Evaluating challenges in existing IDS datasets
Analyzing ML/DL techniques for cybersecurity threat detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Utilizes machine learning for intrusion detection
Reviews multiple datasets like KDDCUP'99
Analyzes classifiers including SVM and ANN
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C V Raman Global University