Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
论文指出编译率作为LLM修复C/C++代码漏洞的指标不可靠,并通过实验验证,提出使用diff_F1作为更有效的初步筛选方法。
📝 Abstract
Large language models (LLMs) are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, and we support this with five controlled experiments over 203 vulnerable functions from Big-Vul, three open-source code LLMs (350M to 6.7B parameters), and three prompting strategies. Compile rate (i) barely responds to an intervention that substantially improves the generated code; (ii) is dominated by evaluation-harness and dataset artifacts rather than model quality, with about 64% of compile failures not attributable to the model, a share that is nearly invariant across models; (iii) shifts by 1.8 to 2.7 times on identical patches under a single compiler-standard flag, with zero regressions; (iv) ranks the three models in the opposite order to reference-similarity metrics; and (v) rewards non-repairs when used as an optimization target, since a compiler-feedback loop raises compile rate while similarity to the human fix falls, with manual inspection finding deletion- and placeholder-style non-repairs among the newly compiling outputs. The natural fallback, whole-function CodeBLEU, also fails: an unchanged copy of the vulnerable input outscores every model. We also examine diff_F1, a change-aware screen that scores only the edited region. It gives exactly zero credit to a no-op and near-zero credit to some, though not all, of the deletion-based gaming patches we observed, while still crediting genuine partial edits, so it may serve as a cheap screen before deeper, execution-based analysis. It is not a repair-quality metric, and we report where it falls short. Our findings argue for change-aware, execution-grounded evaluation of LLM-based vulnerability repair.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Code Vulnerability Repair
Compile Rate
Evaluation Metrics
Change-Aware
Innovation

Methods, ideas, or system contributions that make the work stand out.

change-aware screen
code vulnerability repair
metrics failure
LLMs
diff_F1
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
O
Om Nepal
School of Computing Sciences and Computer Engineering, The University of Southern Mississippi, Hattiesburg, MS 39406, USA
S
Sushant Aryal
School of Computing Sciences and Computer Engineering, The University of Southern Mississippi, Hattiesburg, MS 39406, USA
O
Oluseyi Olukola
School of Computing Sciences and Computer Engineering, The University of Southern Mississippi, Hattiesburg, MS 39406, USA
Nick Rahimi
Nick Rahimi
Associate Professor, University of Southern Mississippi
CybersecurityTrustworthy AIDistributed SystemsP2P Network