🤖 AI Summary
This study addresses the security risks arising from untrusted context injection in large code models, which may lead to the generation of unsafe code. By simulating attack scenarios where malicious instructions are embedded within comments and employing multi-model comparative evaluation alongside equivalence testing techniques, this work systematically investigates the susceptibility of open-source models to web application vulnerabilities. The findings reveal that instruction tuning offers limited mitigation against such attacks and that vulnerability rates are independent of model scale. Under adversarial conditions, the generation rate of high-severity vulnerabilities reaches 77.4%–92.3%, demonstrating that existing screening mechanisms cannot fully eliminate these risks. Ultimately, this research exposes critical security blind spots in code large language models and provides essential empirical evidence for developing robust defense mechanisms.
📝 Abstract
Code large language model (Code LLM) assistants generate code from heterogeneous development contexts, including open files, imported modules, pasted snippets, and comments, much of which may originate from untrusted sources. We investigate whether insecure instructions embedded in such contexts can steer Code LLMs toward vulnerable code without access to model weights or training data. We evaluate ten open-weight Code LLMs spanning 3B--13B parameters, including four base and six instruction-tuned models, across ten web-application weakness classes. We compare completion tasks containing insecure instructions embedded as code comments with benign tasks without malicious instructions. Attack-condition completions contained a medium-or-higher weakness in {\bf 77.4--92.3}\% of cases, compared with {\bf 1.7--5.1}\% in the benign condition. Base and instruction-tuned models averaged 86.5\% and 84.5\% vulnerable outputs, respectively; equivalence testing and three matched model pairs indicated reductions of at most 8.1\% after instruction tuning. Susceptibility showed no clear association with model scale or specialization. Among vulnerable attack outputs, 86.2--91.0\% were rated high or critical, and the effect persisted without the pattern-based detector. Post-generation screening reduced but did not eliminate the risk, the strongest screen leaving roughly one-third undetected. These findings identify inference-time context injection as a substantial attack surface and motivate provenance-aware training objectives.