Know the Normal, Track the Attack: Context-Grounded and Stateful LLM Investigation over System Provenance

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misclassification issues in provenance-based intrusion detection caused by large language models lacking deployment-specific knowledge and insufficient cross-evidence state validation. We propose ANCHOR, a system integrating evidence curation with context-aware reasoning. It introduces a novel anomalous window linking mechanism based on rare relation-role patterns and designs a confidence-gated attack tracking cache to preserve investigative continuity. By combining causality-structure-preserving evidence queue management with kill-chain phase mapping, ANCHOR reconstructs attack narratives from both deployment and case perspectives. Experimental evaluations on DARPA datasets demonstrate that the system significantly improves Indicator of Compromise (IoC) recovery rates and attribution accuracy, enabling continuous round-the-clock audit processing at dollar-level costs.
📝 Abstract
Provenance-based intrusion detection systems (PIDSs) identify suspicious activity in audit streams, but their outputs remain difficult to turn into coherent attack narratives. Direct LLM analyses of local anomalous subgraphs lack deployment-specific normal-behavior knowledge and validated attack state across evidence fragments. This can cause unsupported attack interpretations of routine activities and incorrect attribution of temporally dispersed evidence to attack stages. We present ANCHOR, an investigation-oriented provenance system that combines evidence curation with context-grounded LLM reasoning. It calibrates anomaly judgments by relation type and links anomalous windows through rare relation-role patterns. The resulting evidence queues preserve causal structure, temporal boundaries, and cross-window continuity. The investigator interprets process-centered evidence using two complementary forms of context. Deployment Context combines environment-specific interaction and object baselines with high-risk security knowledge. Case Context uses a confidence-gated Attack-Tracking Cache to maintain investigation state across windows. Correlating current evidence with high-confidence prior findings, ANCHOR incrementally reconstructs attack narratives organized by kill-chain stages. We evaluate ANCHOR on six DARPA Transparent Computing E3/E5 datasets across three operating systems. Controlled evidence-level and end-to-end comparisons show improved overall IoC recovery and attack-stage attribution over state-of-the-art provenance-based baselines. These gains persist under a fixed LLM backbone in our evaluation. ANCHOR processes a full audit day at dollar-level API cost, supporting practical, context-grounded investigation across windows.
Problem

Research questions and friction points this paper is trying to address.

Provenance-based Intrusion Detection
Attack Narrative Reconstruction
Large Language Models
System Provenance
Anomaly Attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Provenance-based Intrusion Detection
Context-Grounded LLM Reasoning
Stateful Attack-Tracking Cache
Evidence Curation
Kill-Chain Reconstruction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lijie Zheng
Xidian University
Ji He
Ji He
Guangzhou Medical University
CT Image ReconstructionDeep Learning
Y
Ying Wang
Xidian University
H
Huang Zhang
Xidian University
Yulong Shen
Yulong Shen
Xidian University
computer security