π€ AI Summary
Public datasets for Network Intrusion Detection Systems (NIDS) suffer from scarcity, unclear applicability, and a lack of standardized quality assessment. Method: We conduct a systematic literature review (SLR), comprehensively analyzing 89 publicly available NIDS datasets across 13 key attributes to enable multidimensional, cross-dataset comparison. Contribution/Results: We propose the first multidimensional evaluation framework for NIDS datasets, including a task-oriented dataset selection guideline, usage best practices, and a critical data quality analysis paradigm. The framework is validated through citation analysis, temporal trend examination, and expert consensus, yielding a reproducible benchmark. Our findings have directly enhanced model robustness and generalization in multiple top-tier conference NIDS studies and are widely adopted as an authoritative reference for dataset selection in the field.
π Abstract
Data-driven cyberthreat detection has become a crucial defense technique in modern cybersecurity. Network defense, supported by Network Intrusion Detection Systems (NIDSs), has also increasingly adopted data-driven approaches, leading to greater reliance on data. Despite its importance, data scarcity has long been recognized as a major obstacle in NIDS research. In response, the community has published many new datasets recently. However, many of them remain largely unknown and unanalyzed, leaving researchers uncertain about their suitability for specific use cases. In this paper, we aim to address this knowledge gap by performing a systematic literature review (SLR) of 89 public datasets for NIDS research. Each dataset is comparatively analyzed across 13 key properties, and its potential applications are outlined. Beyond the review, we also discuss domain-specific challenges and common data limitations to facilitate a critical view on data quality. To aid in data selection, we conduct a dataset popularity analysis in contemporary state-of-the-art NIDS research. Furthermore, the paper presents best practices for dataset selection, generation, and usage. By providing a comprehensive overview of the domain and its data, this work aims to guide future research toward improving data quality and the robustness of NIDS solutions.