Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of existing security inspection frameworks for AI agent skills to malicious evasion. We propose Pretext, an attack framework that leverages white-box large language models to iteratively generate adversarial malicious skills. Its core innovation lies in transforming attack payloads into natural language with decomposed instructions and integrating context injection techniques, thereby simultaneously defeating both static analysis and semantic discrimination mechanisms. Experimental results demonstrate that Pretext achieves bypass rates of up to 97% and 77% against frozen and adaptive detectors, respectively. These findings expose critical fragilities in current agent defense architectures and provide essential insights for developing more robust security evaluation mechanisms.
📝 Abstract
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97\% and 77\% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.
Problem

Research questions and friction points this paper is trying to address.

AI agents
malicious skill detection
evasion attack
skill marketplace
security
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Attack
Malicious Skill Evasion
White-box LLM Attacker
Payload Obfuscation
AI Agent Security
🔎 Similar Papers
💼 Related Jobs
No related jobs found.