vulnerability dataset curation

Designs and builds curated datasets of software vulnerabilities by collecting and consolidating disclosed bugs and reports, mapping and labeling each record to a vulnerability taxonomy, and recording provenance and metadata. Validates and audits annotations, removes duplicates, ensures coverage and representativeness, and manages versioning, licensing, and quality controls for reliable reuse.

vulnerabilitydatasetcuration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$208K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations in evaluating automated vulnerability detection tools, which stem from heterogeneous vulnerability data sources, inconsistent identifiers, and ambiguous version ranges. Leveraging the Open Source Vulnerabilities (OSV) database, we construct a standardized, cross-ecosystem benchmark dataset through precise version mapping and systematic data curation. We propose a reproducible methodology for dataset construction and release an open-source toolkit that enables on-demand reconstruction of historical snapshots, significantly enhancing the transparency and reproducibility of evaluations. Experimental results reveal systematic performance disparities among widely used vulnerability detection tools, underscoring the critical role of high-quality benchmarks in the rigorous assessment of security analysis tools.

evaluationground truthreproducibility

From Bugs to Benchmarks: A Comprehensive Survey of Software Defect Datasets

Apr 24, 2025
HZ
Hao-Nan Zhu
🏛️ University of California, Davis | University of Stuttgart

Large-scale, heterogeneous software defect datasets hinder efficient navigation and reuse by researchers. This paper systematically surveys 132 publicly available defect datasets and proposes a multidimensional evaluation framework—covering domain coverage, defect types, programming language distribution, construction methodologies, accessibility, and citation contexts—to achieve the first standardized metadata harmonization and empirical usability validation across a hundred-plus datasets. Through bibliometric analysis, citation network mapping, and cross-dimensional clustering, we identify test generation and automated program repair as the most widely supported application domains, while critical defect categories—including concurrency and security vulnerabilities—remain severely underrepresented. We further propose actionable dataset curation guidelines and a reusable assessment template, diagnosing pervasive issues such as incomplete coverage, disorganized structure, and outdated maintenance. Our work establishes a robust, empirically grounded data foundation for software defect detection, localization, repair, and AI-driven development.

Evaluating dataset scope, construction, availability, and usabilityIdentifying future research opportunities in defect dataset improvementSurveying 132 software defect datasets for comprehensive analysis

R+R: Security Vulnerability Dataset Quality Is Critical

Mar 09, 2025
AS
Anurag Swarnim Yadav
🏛️ University Of Florida

This paper identifies severe data quality issues in mainstream datasets for vulnerability detection and repair—namely, high duplication rates, pervasive label errors (56%), and incomplete samples (44%)—leading to substantial overestimation of model performance (e.g., VulRepair’s true accuracy drops from the reported 44% to 9–13%). Method: The authors introduce the first systematic quantification of these three quality dimensions and propose a novel paradigm integrating automated deduplication, human-audited label verification, and CWE-guided completeness assessment, alongside constructing a high-quality pretraining corpus. Contribution/Results: Only 31% of samples meet the proposed high-quality standard; leveraging deduplicated bug-fix corpora via transfer learning significantly improves models’ real-world performance. This work establishes the first benchmark for data quality evaluation in vulnerability research and provides a concrete, actionable pathway for dataset curation and model validation.

High duplication rates in vulnerability datasets affect model accuracy.Improved dataset quality enhances Large Language Models' vulnerability detection performance.Incorrect and incomplete labels in datasets lead to unreliable results.

Open-source vulnerability patch datasets suffer from inaccurate labeling and critical sample omissions, severely degrading the performance of downstream security analysis models. To address this, we propose an uncertainty quantification (UQ)-guided, utility-oriented data cleaning framework—the first to integrate model ensembling and heteroscedastic uncertainty modeling for patch quality assessment. Our method combines Monte Carlo Dropout with confidence estimation to automatically identify high-utility patches and filter low-quality samples. This establishes a novel UQ-driven data purification paradigm that preserves dataset representativeness while significantly improving vulnerability prediction accuracy, reducing training time, and lowering computational energy consumption. Extensive experiments demonstrate the effectiveness, scalability, and practical utility of UQ-informed patch selection across diverse vulnerability detection tasks.

Addressing inaccuracies in software vulnerability patch datasetsImproving data quality for machine learning security applicationsReducing negative impact on software security quality assurance

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing vulnerability detection benchmarks, which are predominantly confined to the function level and thus fail to capture cross-procedural vulnerabilities prevalent in real-world scenarios, while manually curated repository-scale datasets suffer from limited scalability. To overcome these challenges, the paper proposes an automated benchmark generation framework that injects realistic vulnerabilities into genuine code repositories and synthesizes reproducible proofs of vulnerability (PoVs), thereby constructing a large-scale, precisely labeled repository-level vulnerability dataset. This framework represents the first scalable approach to automatically generating repository-level vulnerability benchmarks. Furthermore, it introduces an adversarial co-evolution mechanism, enabling dynamic interaction between vulnerability injection and detection agents under realistic constraints, which significantly enhances the robustness and practical utility of vulnerability detection models.

automated dataset generationrealistic vulnerabilityrepository-level benchmark

This work addresses the longstanding challenge in open-source vulnerability datasets—namely, the difficulty of simultaneously achieving reproducibility, scale, and diversity, with reproducibility often compromised to the detriment of automated security research. The authors propose a systematic reproduction framework that integrates version control, automated build and trigger mechanisms, and automatic patch identification. Applying this framework to OSS-Fuzz, they construct ARVO, the first large-scale, reproducible vulnerability dataset comprising over 6,100 real-world vulnerabilities across 311 projects. ARVO enables consistent rebuilds, interactive triggering, and precise patch localization with 89.4% accuracy, achieving an overall reproduction rate of 81%. This significantly enhances the practical utility and research scalability of vulnerability data for security analysis.

bug reproductionopen-source softwarereproducibility

This work addresses the challenges posed by heterogeneous formats and incomplete information in database management system (DBMS) bug reports, which hinder their effective conversion into high-quality test cases. To overcome this, the authors propose BugForge, a novel framework that leverages syntax-aware processing and input-adaptive extraction to construct the first standardized DBMS bug repository from raw proof-of-concept (PoC) artifacts. Integrated with a semantics-guided test case adaptation mechanism, BugForge automatically generates reproducible, high-value test cases. The approach supports fuzz testing, regression testing, and cross-DBMS bug discovery. Evaluated on PostgreSQL, MySQL, MariaDB, and MonetDB using 37,632 historical bug reports, BugForge successfully uncovered 35 new bugs—22 of which have been confirmed—demonstrating significant improvements in both defect detection efficiency and coverage.

automated testingbug report heterogeneitybug repository

This work addresses the lack of security verification datasets that jointly cover requirements, architecture, and code artifacts with fine-grained security annotations. To bridge this gap, the authors introduce EVerest, a novel dataset derived from an open-source electric vehicle charging station software stack, which integrates security requirements, architectural models, source code, and natural language documentation. Through manual extraction, fine-grained annotation, traceability link establishment, and security categorization, EVerest provides the first end-to-end security labeling across these three artifact layers. The dataset comprises 84 security requirements and 1,445 annotated security elements, along with complete supporting artifacts. It has already been applied to identify and remediate real-world CWE vulnerabilities and supports research in security requirement classification, architectural traceability, and code-level verification.

end-to-end securityfine-grained security labelsmulti-artifact dataset

This study addresses the critical issue of inconsistent results from open-source software vulnerability scanners, which significantly hinders informed supply chain security decisions. The work proposes a novel conceptual framework that characterizes the information flows and root causes of inconsistency within the open-source vulnerability ecosystem, modeling vulnerability management as a distributed information transformation process encompassing creation, standardization, enrichment, and contextual interpretation. By integrating multiple vulnerability data standards and real-world case studies, the analysis systematically identifies four core challenges—identity modeling, version semantics, temporal evolution, and contextual assessment—that underlie result discrepancies. This framework establishes a theoretical foundation and offers practical guidance for reproducible evaluation, accurate interpretation of scanner outputs, and dynamic vulnerability knowledge management.

inconsistent findingsopen-source vulnerability ecosystemsoftware supply chain security

Hot Scholars

TY

Terry Yue Zhuo

Researcher
Large Language ModelsCode GenerationAI4SECybersecurity
ZX

Zhenchang Xing

Senior Principal Research Scientist, CSIRO's Data61 & Australian National University
software engineeringhuman-computer-interactionresponsible AI
JZ

Jiayuan Zhou

Principal Researcher, Waterloo Research Centre, Huawei Canada
OSS VulnerabilitiesCrowdsourced Software EngineeringMining Software RepositoriesEmpirical
MS

Miroslaw Staron

Software engineering, University of Gothenburg
Software engineeringmetricsisodependability
SB

Srijita Basu

Postdoc Researcher, Chalmers University of Technology & University of Gothenburg
Cloud ComputingSecurityBlockchainSoftware-defined Networking