🤖 AI Summary
This study addresses the challenge that compact large language models struggle to effectively audit malicious behaviors concealed within agent skills under resource-constrained scenarios. To overcome this limitation, we propose a novel evidence-guided auditing framework that extracts security-relevant behavioral evidence and infers functional intent contexts, thereby guiding compact models toward precise maliciousness assessment despite their inherent constraints in complex reasoning. Extensive experiments demonstrate that this mechanism significantly enhances the detection performance of various compact models, surpassing existing baselines while generalizing effectively to real-world malicious skills. Furthermore, the proposed approach maintains low inference latency, offering an efficient solution for lightweight agent security auditing.
📝 Abstract
As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities. However, the growing adoption of third-party Skills introduces a new supply-chain attack surface. Malicious Skills can embed harmful behaviors that abuse agent privileges and compromise the agent execution environment or accessible resources. Although recent LLM-based malicious Skill auditing approaches have achieved promising performance, they often rely on capable commercial LLMs. How to achieve effective auditing with compact, locally deployable LLMs in security-sensitive and resource-constrained settings remains largely unexplored. Our investigation reveals that compact LLMs struggle to identify malicious behaviors hidden in complex Skill packages. This difficulty arises from both the implicit nature of such behaviors and the limited reasoning capacity of compact LLMs. To address these challenges, we propose SKILLLITE, an evidence-guided agentic framework for malicious Skill detection. SKILLLITE effectively extracts security-relevant behaviors and infers the intended functionality from complex Skill packages. It then employs a compact LLM to assess the maliciousness of the Skill based on the observed behaviors and their functional context. Experiments show that SKILLLITE improves malicious Skill detection across different compact LLM backbones and outperforms existing representative auditing baselines. Its effectiveness generalizes to behaviorally confirmed in-the-wild malicious Skills. Meanwhile, SKILLLITE maintains a low inference latency, supporting its practical deployment.