Chaining Skills to Hijack LLM Agents

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a vulnerability in skill chains of LLM agents, where malicious upstream records can mislead downstream execution into performing attacker-specified actions. To investigate this, the authors propose APEX, a framework that constructs adversarial skill chains through adversarial prompting and multi-skill sequence orchestration. By exploiting task progress records to propagate false authorizations across skills, APEX effectively hijacks agent behavior, revealing that decomposed skills are substantially more susceptible to such attacks than consolidated ones. Furthermore, this work introduces SkillsBench, a dedicated benchmark for evaluating these vulnerabilities. Experimental results demonstrate that APEX achieves an 84.3% targeted attack success rate on GPT-5.4. Notably, existing defense mechanisms significantly degrade normal task performance when attempting to mitigate such attacks, highlighting a critical gap in current agent security paradigms.
📝 Abstract
LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next. Because skills may come from open-source repositories, this handoff can also carry attacker-controlled claims into later decisions. In this paper, we introduce APEX, which constructs and refines adversarial skill chains tailored to a user task and an attacker-selected action. The key insight is that an agent-written record of genuine task progress can carry a false claim of user approval across skills: an upstream skill induces the agent to create the record, and a downstream skill uses it to direct the attacker-selected action. Across four targeted-action families and six models on SkillsBench, the chains induce the selected action in 512 of 690 attempts (74.2%). On GPT-5.4, the full chain succeeds in 84.3% of attempts, compared with 17.4% when the workflow is merged into one skill. We further evaluate a prompting defense that asks the agent to check skill-produced files against the original request. On GPT-5.4, it lowers targeted-action success from 84.3% to 59.1%, while the verifier test-pass rate across 72 benign native-skill tasks falls from 86.7% to 56.3%. These results highlight the need for defenses that prevent attacker-directed actions while preserving legitimate task performance.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
skill chaining
adversarial attack
agent hijacking
AI security
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Skill Chains
LLM Agent Hijacking
Cross-Skill Information Handoff
Prompt Injection
Defense Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.