network traffic dataset curation

Designs and builds benchmark network-traffic datasets — including DoH/3 collections — by collecting and labeling real-world packet- and flow-level traces, parsing raw captures, extracting flow-level, topological and temporal features, and applying preprocessing steps such as filtering, anonymization, aggregation, and normalization. Prepares dataset artifacts and evaluation metadata (e.g., closed/open-world splits, domain-resolution labels, and documentation) for reproducible experiments and public release.

networktrafficdatasetcuration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes an end-to-end, reproducible supervised traffic flow classification framework that addresses the limitations of traditional port- or payload-based methods in the face of encrypted and increasingly diverse network traffic. The framework integrates practical considerations from real-world measurements, incorporating flow-based feature extraction, time-aware data splitting, leakage-proof experimental design, and interpretability analysis to mitigate common methodological pitfalls. Accompanied by an open-source Jupyter Notebook implementation, it provides a complete pipeline—from traffic capture and dataset construction to model training, evaluation, and deployment. Empirical validation on real-world encrypted traffic demonstrates the approach’s effectiveness, robustness, and practical deployability.

encrypted trafficflow-based classificationmachine learning

Technical Report: Generating the WEB-IDS23 Dataset

Feb 06, 2025
EL
Eric Lanfer
🏛️ Osnabrück University

Existing NIDS evaluations suffer from coarse-grained and outdated labels, limited scale, obsolete attack types, and insufficient coverage of modern Web attacks—leading to model overfitting and poor generalization. To address these limitations, this paper introduces WEB-IDS23, a novel dataset specifically designed for Web attack detection. It features a first-of-its-kind modular traffic generator enabling multi-protocol simulation, randomized modeling, and co-synthesis of benign and malicious flows. The dataset provides 82 flow-level features and 21 fine-grained attack classes. Leveraging protocol-aware simulation, stochastic mutation, and pairing with real-world traffic traces, it synthesizes over 12 million labeled samples comprehensively covering prevalent Web attacks (e.g., SQLi, XSS, RCE, path traversal). Empirical evaluation demonstrates that WEB-IDS23 significantly enhances NIDS model representation learning, cross-scenario generalization, and assessment reliability.

Address overfitting with fine-grained labelsDevelop dataset for accurate NIDS evaluationSimulate diverse benign and malicious traffic

Mapping The Invisible Internet: Framework and Dataset

Jun 22, 2025
SA
Siddique Abubakr Muntaka
🏛️ University of Cincinnati | Garden City University College | Kwame Nkrumah University of Science and Technology

Prior research on I2P has predominantly focused on its application layer (e.g., dark web services), leaving a critical gap in systematic, empirical analysis of its network-layer architecture and publicly available measurement datasets. Method: This paper presents the first large-scale network-layer measurement of I2P, introducing SWARM-I2P—a distributed probing framework that integrates dynamic port mapping, netDb parsing, console querying, and passive traffic monitoring. Contribution/Results: We collect and analyze data from over 50,000 I2P routers—including 2,077 FastSet nodes and 2,331 high-capacity routers—along with 4.22 million connection records and over one million packets. The dataset characterizes geolocation (3,444 nodes across 92 countries), bandwidth, latency, uptime, and traffic patterns. To our knowledge, this is the first empirically derived, publicly documented I2P network-layer dataset, enabling rigorous tunnel optimization, resilience assessment, and adversarial modeling—thereby bridging a fundamental gap in anonymous network infrastructure research.

Collecting node and traffic data for anonymity network analysisEnabling resilience studies and adversarial modeling in I2PMapping I2P network layer lacking prior research focus

One task to rule them all: A closer look at traffic classification generalizability

Jul 08, 2025
EA
Elham Akbari
🏛️ University of Waterloo | Orange Labs

Existing traffic classification and website fingerprinting models exhibit severely limited generalization under distribution shift, with performance heavily dependent on specific datasets and environmental assumptions. Method: We propose the first cross-network evaluation framework explicitly targeting real-world network distribution shift—not concept drift—leveraging two large-scale, real-world TLS datasets to conduct cross-domain service identification experiments under a future scenario where Server Name Indication (SNI) is hidden. Results: Even with abundant labeled data, state-of-the-art models achieve only 30–40% accuracy; remarkably, a simple 1-NN classifier performs comparably, challenging the prevailing consensus on the superiority of complex models. Our core contributions are: (i) identifying distribution shift—not concept drift—as the fundamental bottleneck to generalization; (ii) establishing a benchmark framework that isolates and eliminates concept drift confounds; and (iii) empirically demonstrating that lightweight methods offer superior robustness for practical deployment.

Assess classifier performance under real-world distribution shiftEvaluate traffic classification generalizability across different contextsIdentify dataset-specific and model-specific performance limitations

Improving the network traffic classification using the Packet Vision approach

Oct 07, 2020
RM
Rodrigo Moreira
🏛️ Universidade Federal de Viçosa | Universidade Federal de Uberlândia

To address the need for application-aware traffic classification in future smartphone networks, this paper proposes Packet Vision: an end-to-end approach that encodes raw network packets—including headers and payloads—into grayscale images, which are directly fed into convolutional neural networks (e.g., AlexNet, ResNet-18, SqueezeNet) for application-layer classification. This work introduces the first privacy-enhancing image-based paradigm, eliminating explicit plaintext feature extraction while preserving both security and classification accuracy. By designing a customized packet-to-image encoding scheme and constructing a multi-class traffic dataset, the method achieves superior performance on four representative application categories, outperforming state-of-the-art approaches with absolute accuracy gains of 5.2%–12.7%. Experimental results demonstrate the framework’s efficiency, generalizability, and practicality for intelligent network management.

Convolutional Neural NetworksNetwork Traffic ClassificationSmartphone Networks

Latest Papers

What's happening recently
View more

This work addresses the longstanding fragmentation of higher-order network datasets across disparate publications and proprietary platforms, which has hindered their discovery, comparison, and reuse. To overcome this challenge, the authors introduce the Aachen Higher-Order Network Repository (AHORN)—the first structured, traceable, and interoperable centralized repository for such data. AHORN standardizes publicly available higher-order network datasets and provides machine-readable metadata, citation guidelines, format validation, version control, and multi-format export capabilities. By doing so, the platform substantially enhances the centralized management of higher-order network data, facilitates cross-platform reuse, and strengthens scientific reproducibility in the field.

data repositorydata reusedataset discovery

To address the inherent trade-off between accuracy and efficiency in traffic classification on programmable data planes, this paper proposes a lightweight temporal pattern matching framework based on Key Segments. We introduce a novel “offline deep mining–online hardware matching” paradigm, which for the first time compresses packet-level temporal features into hardware-acceleratable key segments. Our approach integrates P4 compiler optimizations, compact hash tables, and SRAM-aware encoding to enable real-time line-rate processing. Experimental evaluation demonstrates that our method achieves a 26.4% improvement in F1-score over statistical baselines and an 18.3% gain over online deep learning approaches, while reducing latency by 13.0% and SRAM footprint by 79.2%. The framework supports full line-rate classification at 100 Gbps on commodity programmable switches.

Achieves high accuracy and efficiency in traffic classificationDeploys packet sequential features via pattern matching on data planesUses Key Segments for efficient matching with reduced resource usage

This study addresses the significant dependence of provenance-based intrusion detection system (PIDS) evaluations on dataset and protocol choices, which often leads to misleading performance comparisons. Conducting a systematic re-evaluation of representative PIDS under a unified temporal split testing protocol and hyperparameter tuning restricted to the validation set—using publicly available datasets that satisfy auditability, labeling, and calibration requirements—the authors find that most reported performance gains stem from lexical novelty in executable names or paths rather than sophisticated provenance modeling. They propose quantifying dataset semantic signal quality via feature completeness and field entropy, which explain model sensitivity to architectural choices. On three of four widely used datasets, a simple allowlist matches or outperforms learning-based methods; only Theia, exhibiting the strongest semantic signals, effectively reveals model advantages in alert prioritization and node recovery.

benchmarkingevaluation protocolforensic investigation

This study investigates the non-deterministic nature of Internet user packet routing paths and the mechanisms by which they are influenced by service providers and IP protocol versions. Leveraging five years of large-scale traceroute measurements spanning six ISP types, twenty autonomous systems, and fourteen countries, the work systematically reveals—across multiple nations, diverse ISP categories, and an extended temporal scale—that user-level paths frequently deviate from geographically shortest routes, often exhibiting significant cross-border detours. The research further demonstrates that transitioning between ISPs or upgrading to IPv6 substantially alters routing policies and end-to-end latency, highlighting the pronounced impact of both administrative and protocol-level factors on path selection in real-world networks.

Internet service providersIP routingIPv4/IPv6

Hot Scholars

SL

Shinan Liu

Assistant Professor, University of Hong Kong
NetworkingSecurityMeasurementMachine Learning Systems
RJ

Raja Jurdak

Professor, Queensland University of Technology (QUT), Australia
Internet of ThingsBlockchainEnergy EfficiencyNetworks
CI

Chadni Islam

Lecturer, Edith Cowan University (ECU)
Cyber SecuritySoftware EngineeringComputer Science
HW

Hassan Wasswa

University of New South Wales (UNSW)
Deep LearningInternet of ThingsCybersecurityComputer Vision
AN

Aziida Nanyonga

University of New South Wales (UNSW), Australia
Artificial IntelligenceMachine learningDeep LearningNatural Language Processing