A Function-level Dataset of Vulnerable and Fixed Source Code in JavaScript and TypeScript

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of high-quality training data for JavaScript and TypeScript vulnerability detection by constructing the JsVul dataset. Methodologically, it designs a language-aware preprocessing pipeline that precisely extracts pre- and post-fix code to isolate security-relevant changes. Furthermore, the approach integrates syntax normalization, multi-stage deduplication, and heuristic labeling techniques to effectively filter noise and retain clean security patches. The resulting dataset is provided in a chronological JSONL format, which significantly enhances model training robustness and detection performance within the JavaScript ecosystem.
📝 Abstract
JavaScript and TypeScript are widely used in modern web development, making their security critical; however, automated vulnerability detection is often constrained by the availability of high-quality training data. Here we present JsVul, a dataset curated from seven major sources. Unlike generic multi-language datasets that may retain noise -- such as minified code and cosmetic edits -- JsVul utilizes a language-specific pipeline. We collected pre-fix and post-fix versions of files around security fixes and, by filtering irrelevant artifacts and applying automated syntax normalization, isolated security-related changes. We ensured data integrity through multi-stage deduplication and heuristic-based labeling. Provided in a time-ordered JSONL format, JsVul supports robust model training in the JavaScript and TypeScript ecosystem and demonstrates the importance of language-aware preprocessing in building vulnerability datasets.
Problem

Research questions and friction points this paper is trying to address.

vulnerability detection
JavaScript
TypeScript
training dataset
source code security
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vulnerability Dataset
Language-specific Pipeline
Syntax Normalization
Automated Deduplication
Heuristic Labeling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.