Beware EviLLM: Enabling Vulnerability Injection via Large Language Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the realistic threat of third-party adversaries maliciously injecting security vulnerabilities through AI-assisted code generation pipelines. To this end, it proposes EviLLM, an attack framework that constructs a practical third-party threat model distinct from unintentional vulnerability introduction or insider threats. Specifically, EviLLM leverages compromised accounts or browser hijacking to intercept and manipulate both large language models and natural language specifications, enabling automated vulnerability injection via existing web-based attack vectors. Experimental results demonstrate that the proposed framework successfully introduces vulnerabilities spanning 13 Common Weakness Enumeration (CWE) categories. Furthermore, a user study confirms that most developers fail to detect these injected security risks. These findings highlight the significant fragility of the software supply chain within AI-assisted programming ecosystems.
📝 Abstract
Advances in large language models (LLMs) have enabled AI-driven code generation from natural language specifications, introducing new attack surfaces for injecting vulnerabilities into software. Prior work has studied this problem only in benign settings where vulnerabilities are introduced inadvertently, or under unconventional threat models where the LLM itself is malicious (backdooring) or the user is the attacker (jailbreaking). In this paper, we study a more realistic threat model: a third-party adversary, with capabilities comparable to existing cybercriminals, compromises the AI code generation pipeline to deliberately introduce vulnerabilities. We call this the EviLLM attack. We have implemented two instances of EviLLM, each of which only requires the underlying LLM to be accessed through a compromised account or browser, and can inject vulnerabilities from 13 CWE classes. As we show in our feasibility study, both attack vectors are already used to implement many existing cyberattacks. Our user study shows that 7 out of 8 and 10 out of 13 participants did not notice the vulnerabilities injected by the two instances of EviLLM, and 13 out of 21 participants"rarely"or"never"considered the risk of an attack like EviLLM.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Vulnerability Injection
AI Code Generation
Threat Model
Software Security
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Vulnerability Injection
EviLLM Attack
Code Generation Security
Threat Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zeezoo Ryu
Georgia Institute of Technology
S
Simon Chung
Georgia Institute of Technology
M
Muhammad Faraz Karim
Georgia Institute of Technology
A
Anna Raymaker
Georgia Institute of Technology
K
Karan Singh Jodha
Georgia Institute of Technology
Y
Yash Chaturvedi
Snowflake
S
Sukarno Mertoguno
Georgia Institute of Technology