CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of evaluation frameworks for LLM-based agents performing APT attack chain provenance tracing in complex security logs by proposing CyberProvenance. The framework constructs a benchmark encompassing single-step and multi-stage attacks, requiring agents to identify evidence, infer processes, and generate provenance graphs containing entities and causal relationships. It introduces a semantic consistency-based multidimensional evaluation method and designs a multi-agent collaboration mechanism integrating evidence accumulation with feedback-driven correction, combined with MITRE ATT&CK mapping and execution verification techniques. Experimental results demonstrate that this approach significantly enhances evidence reasoning and complete attack chain reconstruction capabilities, effectively overcoming the limitations of existing systems in long-context provenance analysis.
📝 Abstract
Large language model agents have demonstrated promising capabilities in cybersecurity tasks, yet their ability to reconstruct complete Advanced Persistent Threat attack campaigns from complex security logs remains largely unexplored. Existing cybersecurity benchmarks for agents mainly focus on vulnerability discovery, exploitation, and security analysis tasks, leaving the evaluation of attack chain provenance under realistic security logs insufficiently studied. To address this gap, we introduce CyberClear, a benchmark for evaluating LLM agents and advanced agent systems on APT attack chain provenance from long-context security logs. CyberClear covers both single-step attacks and multi-stage attack chains, requiring agents to identify attack evidence, infer attack progression, and generate provenance graphs containing entities, causal relationships, MITRE ATT&CK techniques, and forensic evidence. To enable comprehensive evaluation, we develop an evaluation method tailored to APT attack chain provenance. Unlike conventional text similarity metrics that focus on surface-level matching, our evaluation examines whether reconstructed graphs preserve the semantics of attack chains across single-step behavior correctness, multi-step behavior identification, temporal and causal consistency, entity and relationship fidelity, and overall attack narrative consistency. Advanced multi-agent systems powered by state-of-the-art LLMs still struggle on CyberClear, motivating us to propose CyberProvenance, an agent cyber harness designed for multi-agents that augments LLM agents with evidence accumulation, execution-based validation, and feedback-guided refinement mechanisms for reliable attack-chain provenance. Extensive evaluations on CyberClear demonstrate the effectiveness of CyberProvenance in improving evidence reasoning, execution-grounded validation, and complete APT attack chain reconstruction.
Problem

Research questions and friction points this paper is trying to address.

LLM Agent
APT Attack Chain
Provenance
Cybersecurity Benchmark
Security Logs
Innovation

Methods, ideas, or system contributions that make the work stand out.

APT Attack Chain Provenance
LLM Agent Benchmark
Multi-Agent System
Provenance Graph Evaluation
Evidence Accumulation
🔎 Similar Papers
No similar papers found.
Q
Qi Chen
School of Cyber Science and Engineering, Southeast University
Fushuo Huo
Fushuo Huo
The Hong Kong Polytechnic University
Large Vision Language ModelMultimodal LearningTrustworthy AI
H
Hangli Shen
School of Cyber Science and Engineering, Southeast University
Jingcai Guo
Jingcai Guo
Hong Kong Polytechnic University
Efficient AIZero-Shot LearningEdge AIMachine Learning
Shuhao Li
Shuhao Li
Fudan University
G
Guang Cheng
School of Cyber Science and Engineering, Southeast University