🤖 AI Summary
Existing benchmarks overlook post-compromise persistence, failing to assess the realistic risk of LLM agents maintaining footholds under system disruptions. This work constructs the first post-exploitation persistence benchmark, reframing the task as an adversarial survival challenge. It introduces a six-layer deterministic scoring framework to decouple vulnerability exploitation from persistence evaluation, integrating multi-host architectures with active defense simulations for automated attack-defense analysis. Experiments reveal that frontier LLMs achieve autonomous persistence success rates of only 27.6%–44.8%, which drop sharply to 5.5%–13.3% when defenses are enabled. This study bridges the gap in evaluating long-term dwell capabilities, exposes emerging cyberattack risks, and delineates the operational boundaries of autonomous agents.
📝 Abstract
While LLM-based attackers exhibit growing proficiency in vulnerability exploitation, most existing cybersecurity benchmarks suffer from single-stage truncation, prematurely terminating evaluation upon initial access. In practice, initial footholds are exceptionally fragile across operational disruptions such as service restarts and host reboots. Whether LLM-based attackers can establish and maintain durable footholds beyond initial compromise remains a central blind spot in cybersecurity evaluation. To bridge this gap, we introduce CyberPersistBench, the first benchmark dedicated to post-compromise installation and persistence. Decoupled from upfront exploitation, CyberPersistBench frames persistence as an adversarial survival task in which agents use native host mechanisms to maintain footholds across staged system disruptions. Deterministic checks support a six-level scoring method (L1--L6) spanning installation and persistence. The benchmark comprises 203 core tasks across seven categories, augmented by multi-host and active defense extensions. Empirical evaluations across five frontier agents show that autonomous persistence remains limited (27.6%--44.8%) and drops further on defense-enabled tasks (5.5%--13.3%); nonetheless, these results reveal an emerging cyberattack risk, underscoring the necessity of benchmarking post-compromise persistence. CyberPersistBench thus establishes a foundational benchmark for post-compromise installation and persistence, delineating the operational boundaries of autonomous cyber agents.