Where Cyber Agents Struggle: Bottleneck Analysis of Multi-Stage LLM Agents

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of multi-stage LLM agents in autonomous adversarial scenarios to cost inefficiency, excessive retries, and evidence misinterpretation, noting that success rate alone fails to reveal their adaptive shortcomings. Based on an orchestration–execution–verification three-tier architecture, this work conducts end-to-end diagnostics on six frontier models within lateral movement scenarios. It introduces a novel subtask-conditioned, cost-aware scoring metric and employs LLM-as-a-Judge to identify planning deficiencies. The findings indicate that credential acquisition and lateral movement constitute core bottlenecks, while verifiers consistently exhibit over-optimistic biases. These results underscore the necessity of comprehensively evaluating both resource efficiency and system adaptability when deploying LLM agents in complex adversarial environments.
📝 Abstract
Multi-stage LLM-based cyber agents may complete attack workflows while remaining brittle, costly, or reliant on incorrect interpretations of execution evidence. Success rates alone obscure inefficiency, adaptation through retries, and recognition of success or failure. We present an end-to-end diagnostic study of an Autonomous Adversary system with orchestrator, executor, and validator LLMs in enterprise-like lateral-movement scenarios. Six frontier models are evaluated across two scenarios and three modes: expert-defined, self-scaffolded, and fully autonomous. We assess validator consistency and evidence grounding; introduce a subtask-conditioned, cost-aware score for abnormal token use, retries, and runtime; and use comparative LLM-as-a-Judge analysis to identify planning deficiencies, including tool misalignment, plan similarity, over-specification, inadequate probing, and weak recovery. Validators are generally relevant and evidence-grounded but often nonspecific and overly optimistic. Bottlenecks cluster in credential and lateral-movement tasks, spread with scenario complexity, and vary more under full autonomy. Reliable evaluation must assess outcomes, evidence interpretation, resource use, and adaptation after failure.
Problem

Research questions and friction points this paper is trying to address.

Multi-stage LLM Agents
Cyber Agents
Bottleneck Analysis
Autonomous Adversary
Evaluation Metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Stage LLM Agents
Bottleneck Analysis
Cost-Aware Scoring
LLM-as-a-Judge
Autonomous Adversary
🔎 Similar Papers