🤖 AI Summary
This study addresses the sharp decline in injection attack success rates caused by routing contention in multi-skill systems. To overcome this limitation, we propose CORSA, a method that introduces router-aware mechanisms into adversarial design for the first time. By leveraging cluster optimization and multi-stage iterative strategies, CORSA synergistically enhances both the retrieval hit rate of malicious skills and end-to-end execution success. Furthermore, we extend the SkillRouter benchmark to evaluate eight categories of malicious payloads. Experimental results demonstrate that CORSA significantly improves attack effectiveness while preserving user utility, exhibiting strong transferability across diverse architectures and models.
📝 Abstract
AI agents increasingly rely on modular third-party "skills" that are dynamically selected by skill routers to execute complex tasks. While recent studies highlight the threat of prompt injections embedded in these skills, existing evaluations often assume settings where the malicious skill is already selected for execution. We show that this assumption can substantially overestimate attack success. In realistic multi-skill environments, injected skills must first compete for retrieval, reducing the effective attack success rate (ASR) of existing injections by 87-97%. To address this limitation, we introduce CORSA (Cluster Optimization for Router-Aware Skill Attacks), a router-aware attack that optimizes skill injections for both retrieval and execution across clusters of related tasks. We evaluate skill injection attacks under router-managed multi-skill settings by extending the benchmark introduced by SkillRouter with eight malicious payload categories. CORSA uses successive optimization stages to first improve retrieval and then optimize end-to-end attack success, while we evaluate user utility and injection naturalism separately. Our experiments show that CORSA substantially improves both retrieval and end-to-end attack success over existing skill injections while preserving user utility, and that the resulting attacks transfer across different router architectures and LLM backbones.