🤖 AI Summary
This study addresses the realistic threat of third-party adversaries maliciously injecting security vulnerabilities through AI-assisted code generation pipelines. To this end, it proposes EviLLM, an attack framework that constructs a practical third-party threat model distinct from unintentional vulnerability introduction or insider threats. Specifically, EviLLM leverages compromised accounts or browser hijacking to intercept and manipulate both large language models and natural language specifications, enabling automated vulnerability injection via existing web-based attack vectors. Experimental results demonstrate that the proposed framework successfully introduces vulnerabilities spanning 13 Common Weakness Enumeration (CWE) categories. Furthermore, a user study confirms that most developers fail to detect these injected security risks. These findings highlight the significant fragility of the software supply chain within AI-assisted programming ecosystems.
📝 Abstract
Advances in large language models (LLMs) have enabled AI-driven code generation from natural language specifications, introducing new attack surfaces for injecting vulnerabilities into software. Prior work has studied this problem only in benign settings where vulnerabilities are introduced inadvertently, or under unconventional threat models where the LLM itself is malicious (backdooring) or the user is the attacker (jailbreaking). In this paper, we study a more realistic threat model: a third-party adversary, with capabilities comparable to existing cybercriminals, compromises the AI code generation pipeline to deliberately introduce vulnerabilities. We call this the EviLLM attack. We have implemented two instances of EviLLM, each of which only requires the underlying LLM to be accessed through a compromised account or browser, and can inject vulnerabilities from 13 CWE classes. As we show in our feasibility study, both attack vectors are already used to implement many existing cyberattacks. Our user study shows that 7 out of 8 and 10 out of 13 participants did not notice the vulnerabilities injected by the two instances of EviLLM, and 13 out of 21 participants"rarely"or"never"considered the risk of an attack like EviLLM.