🤖 AI Summary
This study addresses the scarcity of high-quality training data for JavaScript and TypeScript vulnerability detection by constructing the JsVul dataset. Methodologically, it designs a language-aware preprocessing pipeline that precisely extracts pre- and post-fix code to isolate security-relevant changes. Furthermore, the approach integrates syntax normalization, multi-stage deduplication, and heuristic labeling techniques to effectively filter noise and retain clean security patches. The resulting dataset is provided in a chronological JSONL format, which significantly enhances model training robustness and detection performance within the JavaScript ecosystem.
📝 Abstract
JavaScript and TypeScript are widely used in modern web development, making their security critical; however, automated vulnerability detection is often constrained by the availability of high-quality training data. Here we present JsVul, a dataset curated from seven major sources. Unlike generic multi-language datasets that may retain noise -- such as minified code and cosmetic edits -- JsVul utilizes a language-specific pipeline. We collected pre-fix and post-fix versions of files around security fixes and, by filtering irrelevant artifacts and applying automated syntax normalization, isolated security-related changes. We ensured data integrity through multi-stage deduplication and heuristic-based labeling. Provided in a time-ordered JSONL format, JsVul supports robust model training in the JavaScript and TypeScript ecosystem and demonstrates the importance of language-aware preprocessing in building vulnerability datasets.