AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks

📅 2026-02-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of large language model (LLM) agents to multi-round adversarial attacks in prolonged, complex interactions—a threat inadequately mitigated by existing single-turn defense mechanisms. We present the first systematic formulation of “long-term attacks” and introduce a comprehensive security evaluation benchmark tailored to this threat model, encompassing five novel attack categories: intent hijacking, toolchain abuse, task injection, and others. The benchmark includes 644 test cases across 28 realistic environments. To enable scalable assessment, we develop an adversarial testing platform that simulates multi-turn user–agent–environment interactions and integrates automated attack generation and evaluation. Empirical results demonstrate that state-of-the-art LLM agents are highly susceptible to these long-term attacks, and conventional defense strategies largely fail, underscoring the urgent need for new security paradigms.

Technology Category

Multiagent Systems: Adversarial AgentsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and Robustness

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
LLM agents are increasingly deployed in long-horizon, complex environments to solve challenging problems, but this expansion exposes them to long-horizon attacks that exploit multi-turn user-agent-environment interactions to achieve objectives infeasible in single-turn settings. To measure agent vulnerabilities to such risks, we present AgentLAB, the first benchmark dedicated to evaluating LLM agent susceptibility to adaptive, long-horizon attacks. Currently, AgentLAB supports five novel attack types including intent hijacking, tool chaining, task injection, objective drifting, and memory poisoning, spanning 28 realistic agentic environments, and 644 security test cases. Leveraging AgentLAB, we evaluate representative LLM agents and find that they remain highly susceptible to long-horizon attacks; moreover, defenses designed for single-turn interactions fail to reliably mitigate long-horizon threats. We anticipate that AgentLAB will serve as a valuable benchmark for tracking progress on securing LLM agents in practical settings. The benchmark is publicly available at https://tanqiujiang.github.io/AgentLAB_main.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
long-horizon attacks
agent security
multi-turn interactions
vulnerability benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

long-horizon attacks
LLM agents
security benchmark
adaptive attacks
agent vulnerability