A Mixture of Linear Corrections Generates Secure Code

📅 2025-07-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) exhibit strong code-generation capabilities but struggle to reliably avoid security vulnerabilities, primarily because existing prompting techniques fail to effectively activate their latent vulnerability-discrimination abilities. This paper proposes MoC (Modulated Correction), a fine-tuning-free, inference-time steering method that identifies security- and vulnerability-sensitive internal representations—specifically in hidden layers—via representation engineering, and dynamically rectifies token probability distributions through a linear mixture mechanism. Experiments on Qwen2.5-Coder-7B demonstrate that MoC simultaneously enhances both security and functionality: secure code generation rate improves by 8.9%, and HumanEval pass@1 increases by 2.1%. The core contribution lies in the first identification and exploitation of LLMs’ intrinsic, pre-trained vulnerability-aware representations, enabling efficient, lightweight, and plug-and-play security augmentation without architectural or training modifications.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision ModelsNatural Language Processing: Safety and Robustness

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSecurity and Privacy: Large-scale security measurementsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
Large language models (LLMs) have become proficient at sophisticated code-generation tasks, yet remain ineffective at reliably detecting or avoiding code vulnerabilities. Does this deficiency stem from insufficient learning about code vulnerabilities, or is it merely a result of ineffective prompting? Using representation engineering techniques, we investigate whether LLMs internally encode the concepts necessary to identify code vulnerabilities. We find that current LLMs encode precise internal representations that distinguish vulnerable from secure code--achieving greater accuracy than standard prompting approaches. Leveraging these vulnerability-sensitive representations, we develop an inference-time steering technique that subtly modulates the model's token-generation probabilities through a mixture of corrections (MoC). Our method effectively guides LLMs to produce less vulnerable code without compromising functionality, demonstrating a practical approach to controlled vulnerability management in generated code. Notably, MoC enhances the security ratio of Qwen2.5-Coder-7B by 8.9%, while simultaneously improving functionality on HumanEval pass@1 by 2.1%.
Problem

Research questions and friction points this paper is trying to address.

Detect and avoid code vulnerabilities in LLM-generated code
Investigate internal representations of code vulnerabilities in LLMs
Enhance code security without compromising functionality via MoC
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses representation engineering to detect vulnerabilities
Applies mixture of corrections for token modulation
Enhances code security without losing functionality