🤖 AI Summary
This work addresses the critical challenges in code vulnerability detection—namely, the scarcity of high-quality labels, substantial noise, and imbalanced data distributions—all of which heavily rely on costly manual annotation. The study presents the first systematic mapping framework for label-efficient vulnerability detection, organizing existing approaches into five paradigm families: weak supervision, self-supervision, transfer learning, and others. It further links these paradigms to diverse code representations, including token-based, graph-based, hybrid, and knowledge-enhanced forms. By introducing a design taxonomy and a constraint-prioritized decision guide, the paper clarifies the applicability and failure modes of each method, exposes key obstacles such as inconsistent evaluation protocols, and establishes a unified perspective for evaluating and selecting techniques under label-scarce conditions.
📝 Abstract
Machine-learning-based code vulnerability detection (CVD) has progressed rapidly, from deep program representations to pretrained code models and LLM-centered pipelines. Yet dependable vulnerability labeling remains expensive, noisy, and uneven across projects, languages, and CWE types, motivating approaches that reduce reliance on human labeling. This survey maps these approaches, synthesizing five paradigm families and the mechanisms they use. It connects mechanisms to token, graph, hybrid, and knowledgebased representations, and consolidates evaluation and reporting axes that limit comparison (label-budget specification, compute/cost assumptions, leakage, and granularity mismatches). A Design Map and constraintfirst Decision Guide distill trade-offs and failure modes for practical method selection.