JevVibe: Efficient Classification-Guided Secure Code Generation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the security vulnerabilities in LLM-generated code and the inefficiency of traditional autoregressive CWE classification. We propose the Jev decision framework, built upon discriminative models, alongside the JevVibe repair agent. Departing from generative paradigms, our approach directly assigns CWE labels to candidate sets through discriminative inference, leveraging these predictions to guide automated secure code repair and thereby establishing an efficient diagnosis-repair closed loop. Evaluated on the CyberSecEval benchmark, the proposed method comprehensively outperforms open-source baselines. Specifically, it achieves a 6.27× improvement in inference speed and a 55.9× reduction in computational cost compared to GPT-5.6-Sol, while increasing the code security pass rate from 63.5% to 70.7%.
📝 Abstract
Large language models can generate functionally correct code that still contains security weaknesses, motivating repair pipelines that first diagnose a weakness type before deciding how to fix it. The Common Weakness Enumeration (CWE) provides a standardized vocabulary for such diagnoses, but asking an autoregressive language model to generate a CWE label and extracting it from the response raises questions about output validity, speed, and cost, as well as accuracy. We evaluate Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, GPT-5.6-Sol, on a controlled 50-way CWE classification task over 1,916 CyberSecEval benchmark examples. Jev outperforms all six open-weight baselines on every classification and ranking metric, while its comparison with GPT-5.6-Sol depends on the metric: GPT-5.6-Sol achieves higher Top-1 accuracy and Macro-F1, whereas Jev achieves higher Top-3 and Top-5 accuracy and a nearly identical MRR, at $6.27\times$ lower median API latency and $55.9\times$ lower estimated API cost. We further build JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code generated by Qwen2.5-Coder-32B-Instruct. With Jev providing the diagnosis, the agent increases the detector-measured security pass rate from 63.5% before repair to 70.7%, compared with 66.1% for LLM-guided repair. These results show that JevVibe is effective at improving the security of generated code, with Jev providing reliable and efficient CWE classification.
Problem

Research questions and friction points this paper is trying to address.

Secure Code Generation
Common Weakness Enumeration
Code Vulnerability Repair
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Secure Code Generation
CWE Classification
Decision Model
Diagnosis-Guided Repair
Efficiency
🔎 Similar Papers