🤖 AI Summary
This study addresses the vulnerability of autonomous coding agents to prompt injection threats when relying on community rule files. Specifically, this work introduces "package hallucination" attacks, wherein malicious prompts injected into benign rule files induce agents to replace legitimate dependencies with attacker-controlled malicious packages. To automate and optimize such attacks, we propose PackHallu, a framework that integrates evolutionary optimization, large language model-guided mutation, and trajectory-level feedback to automatically generate and iteratively refine effective yet stealthy injection prompts. Extensive experiments demonstrate that PackHallu achieves high attack success rates and strong transferability across multiple benchmarks, foundation models, and agent frameworks. These findings expose critical security vulnerabilities in current automated coding systems, underscoring the urgent need for robust defense mechanisms against dependency-level prompt injection attacks.
📝 Abstract
Modern agentic coding frameworks increasingly rely on community-shared rule files (e.g., AGENTS.md or .cursorrules) to guide autonomous code generation, yet the security risks of this pipeline remain underexplored. To bridge this gap, we introduce the package hallucination attack, where an attacker injects malicious prompts into benign rule files to induce coding agents to replace legitimate dependencies with attacker-controlled packages. To obtain effective malicious prompts injected into rule files, we propose PackHallu, an evolutionary optimization framework that iteratively rewrites these injected prompts using trajectory-level feedback and LLM-guided mutations. Evaluations across multiple benchmarks, LLMs, and agent frameworks show that PackHallu achieves high attack success rates and strong transferability across diverse models and agent combinations. Our findings demonstrate that coding agents are vulnerable to package hallucination attacks, highlighting the urgent need for stronger security safeguards in autonomous coding systems.