Score
Designs and builds synthetic network traffic datasets and simulators that produce packet- and flow-level records—including benign, malicious, and adversarially perturbed samples—to augment training data, impute or reconstruct incomplete traces, and generate labeled data for benchmarking. Creates generative models and traffic‑modeling pipelines to simulate periodic and evolving patterns, varying loads, and evasive attack behavior for stress‑testing, evaluating, and hardening intrusion detection systems and other network analytics.
This study addresses the challenges confronting intrusion detection systems (IDS), including dynamically evolving attacks, data scarcity, class imbalance, and the difficulty of centralized training under strict privacy constraints. It presents a systematic review of the integration of generative artificial intelligence—encompassing generative adversarial networks (GANs), diffusion models, and large language models—with federated learning in IDS applications. The work covers key areas such as anomaly detection, synthetic data generation, data augmentation, and privacy-preserving distributed training. For the first time, it offers a structured synthesis of recent advances at the intersection of these two technological paradigms and outlines promising future directions, including domain-specific large language models and federated benchmarking frameworks, thereby establishing a novel paradigm for privacy-sensitive, distributed cybersecurity solutions.
Synthetic network traffic data is critical for security testing and model training, yet existing generation methods exhibit inconsistent performance in statistical fidelity, classification utility, and class balance. This paper systematically evaluates 12 generative techniques—including statistical approaches (e.g., SMOTE), classical AI models, and modern generative AI (e.g., CTGAN, CopulaGAN, diffusion models)—across NSL-KDD and CIC-IDS2017 datasets, using unified metrics for fidelity, downstream classification performance, class balance, and scalability. We introduce the first multi-dimensional benchmarking framework tailored to network traffic synthesis. Results show CTGAN and CopulaGAN achieve the best trade-off between fidelity and classification utility; statistical methods yield superior class balance but limited modeling capacity; diffusion models, while promising, suffer from prohibitive computational overhead, hindering practical scalability. The study provides empirically grounded guidance for selecting synthetic data generation methods in cybersecurity applications.
Existing NIDS evaluations suffer from coarse-grained and outdated labels, limited scale, obsolete attack types, and insufficient coverage of modern Web attacks—leading to model overfitting and poor generalization. To address these limitations, this paper introduces WEB-IDS23, a novel dataset specifically designed for Web attack detection. It features a first-of-its-kind modular traffic generator enabling multi-protocol simulation, randomized modeling, and co-synthesis of benign and malicious flows. The dataset provides 82 flow-level features and 21 fine-grained attack classes. Leveraging protocol-aware simulation, stochastic mutation, and pairing with real-world traffic traces, it synthesizes over 12 million labeled samples comprehensively covering prevalent Web attacks (e.g., SQLi, XSS, RCE, path traversal). Empirical evaluation demonstrates that WEB-IDS23 significantly enhances NIDS model representation learning, cross-scenario generalization, and assessment reliability.
This study addresses the critical challenges in encrypted traffic analysis—namely, data invisibility and severe class imbalance due to the scarcity of anomalous samples. To tackle these issues, this work introduces, for the first time, a systematic application of generative artificial intelligence to the domain, proposing a synthetic data generation method that preserves the statistical properties and feature correlations of original traffic. The approach integrates feature analysis, clustering guidance, and generative modeling to collaboratively construct a balanced dataset. Experimental results demonstrate that classifiers trained on the synthesized data achieve 93% of the performance of models trained on real data, substantially enhancing anomaly detection capabilities. The implementation code has been made publicly available.
To address key bottlenecks in network traffic analysis—namely poor few-shot adaptability, strong label dependency, and weak cross-domain generalization—this paper proposes the first generative pre-training framework specifically designed for NetFlow data. Our method employs a Transformer-based architecture with a self-supervised generative pre-training objective, enabling unified representation learning and large-scale unsupervised modeling of network flow features. A lightweight fine-tuning mechanism is introduced to rapidly adapt the pre-trained model to diverse downstream tasks, including classification, congestion prediction, and DDoS detection. Evaluated on real-world DDoS detection, the approach achieves 92.4% accuracy using only limited labeled data—outperforming supervised baselines by 12.7%. Moreover, the pre-trained model demonstrates strong transferability across heterogeneous network environments. This work establishes the first unified pre-training paradigm for NetFlow representation learning, advancing foundational methodology for traffic analytics.
Existing network simulation models neglect application-layer behavior, leading to traffic distortion and hindering robustness evaluation of monitoring and anomaly detection systems. To address this, we propose the first framework that models application-layer behavior as learnable and composable probabilistic processes. Specifically, it estimates probability density functions from real-world traffic traces and employs behavioral pattern convolution to generate dynamic, scalable, and realistic traffic. We further design a lightweight simulation engine supporting coexistence of multiple applications on a single machine and real-time behavioral modulation. Experimental results demonstrate that our approach significantly improves test coverage while achieving traffic distributions closely aligned with real-world scenarios. The open-source implementation has been validated on large-scale production networks.
This study addresses the limitations of traditional intrusion detection systems in digital forensics, particularly their lack of traceability, reproducibility, and courtroom admissibility, as well as their inability to provide instance-level explanations. To overcome these challenges, the authors propose a forensically compliant intrusion detection framework aligned with judicial standards such as ISO/IEC 27037, which strictly segregates original evidence from analytical artifacts. The approach leverages Synthetic Data Vault (SDV) and CTGAN to generate synthetic network traffic for training an XGBoost classifier, while SHAP TreeExplainer enables instance-level attribution of attack behaviors. Experimental results demonstrate that synthetic data achieves a TSTR F1-macro score of 0.96 on CICIDS2017, and cross-dataset validation confirms that approximately 30 key features suffice to maintain detection performance. Critically, SHAP attributions exhibit high consistency between real and synthetic data, effectively preserving attack fingerprints while ensuring both high accuracy and forensic compliance.
This work addresses the limitations of existing cybersecurity datasets, which are predominantly static and ill-suited for enabling controllable replay and traceability in heterogeneous, multi-protocol environments. To overcome this, the authors propose a scenario-oriented, container-native testing platform that leverages declarative configuration to parameterize the generation of both adversarial and benign network traffic, log collection, and dataset integration. The platform encapsulates 60 attack scenarios, nine target services, and benign traffic generators within single-purpose containers and integrates them into an automated pipeline for feature extraction and experimental execution. Designed with reproducibility, auditability, and extensibility in mind, the framework significantly reduces operational bias and supports fully traceable, reproducible experiments in complex settings such as IoT and IIoT networks.
This study addresses the limitations of existing IoT intrusion detection datasets, which commonly suffer from fixed attack categories and extreme class imbalance, as well as the inability of current generative models to guarantee the physical validity of synthesized packets. To overcome these challenges, the paper proposes two novel synthesis approaches that embed hard validity constraints directly into the generation process: a statistical learning method based on PCA and dual anomaly detection boundaries, and a genetic algorithm that formulates data generation as a multi-objective optimization problem. Innovatively integrating dual anomaly gating, feature-range clamping, and an independent validation mechanism, both methods significantly enhance the fidelity and validity of synthetic data. Evaluated on the ACI IoT 2023 dataset, they achieve average anomaly rates of 1.20% (at 1,091 pkt/s) and 0.62% (at 5.7 pkt/s), respectively, and successfully expand ARP spoofing samples to 1,000 instances—a 200-fold increase.
This study addresses the challenges of network traffic classification under scarce labeled data and stringent privacy constraints, where conventional generative approaches struggle to balance temporal modeling fidelity with computational efficiency. For the first time, it systematically evaluates lightweight generative AI models—including Transformers, state space models, and diffusion models—for synthetic traffic generation, proposing an efficient and privacy-preserving synthesis framework. The generated traffic accurately preserves both static and dynamic temporal characteristics. Notably, classifiers trained exclusively on synthetic data achieve an F1-score of 87% on real-world traffic. Furthermore, in low-data regimes, the proposed data augmentation strategy improves classification performance by up to 40%, substantially narrowing the gap with models trained on full datasets.
This work addresses key challenges in network intrusion detection—namely data scarcity, privacy sensitivity, and insufficient model robustness—by introducing a novel multimodal dataset that unifies traffic, payload, and temporal contextual features into a cohesive representation space. To enhance data availability while preserving privacy, the study proposes a synthetic data generation approach that integrates adversarial generative models with the Synthetic Data Vault (SDV) framework. The fidelity, utility, and privacy-preserving properties of the generated data are rigorously validated through f-divergence metrics, distinguishability tests, TRTS/TSTR evaluations, and non-parametric statistical analyses. Experimental results demonstrate that the proposed method significantly improves the accuracy and generalization capability of intrusion detection models, thereby establishing a high-quality, reproducible foundation for cybersecurity research and evaluation.