🤖 AI Summary
This study addresses the issue of function-level label noise in vulnerability detection datasets caused by commit-level associations. To this end, it proposes the first automated label auditing framework grounded in executable evidence. Specifically, the method employs large language model agents to orchestrate dynamic analysis tools, advancing vulnerability label verification from static inference to a dynamic validation paradigm through trigger experiment construction, runtime behavior comparison, and automated input refinement. Experimental results demonstrate that the proposed framework successfully rectifies nearly 20% of erroneous labels, achieving an accuracy exceeding 90% as confirmed by blind review. These corrections substantially enhance both the evaluation reliability and training efficacy of downstream models.
📝 Abstract
Reliable learning-based vulnerability detection requires high-quality labels, yet datasets built from vulnerability-fixing commits may label functions as vulnerable simply because they were changed by a security patch. We present VulValidate, a framework that uses LLM agents to coordinate dynamic analysis tools and construct vulnerability-triggering experiments from runtime feedback. Given a labeled function and its fixing patch, VulValidate reconstructs vulnerable and fixed revisions, selects suitable tools and execution paths, refines triggering inputs, and compares runtime behavior to assess function-level attribution. We audit all 35,849 instances originally labeled vulnerable in BigVul, PrimeVul, and DiverseVul. We confirm 20,510 (57.2%), correct 6,819 labels (19.0%), leave 7,981 attacked but undecided (22.3%), and cannot successfully measure 539 (1.5%). After conflict resolution and byte-exact deduplication, the corrected release contains 15,890 distinct confirmed vulnerable function bodies. In a blinded review of 581 sampled decisions, expert consensus supports 90.0%--92.0% of confirmations and 92.6%--99.0% of label corrections. With model parameters fixed, corrected evaluation lowers F1 for all five tested detectors on both BigVul and DiverseVul; retraining with corrected labels improves F1 for four of five detectors on each dataset. We also release a reusable VulValidate skill, corrected datasets, and reproducible evidence for future vulnerability-detection research.