How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant dependence of provenance-based intrusion detection system (PIDS) evaluations on dataset and protocol choices, which often leads to misleading performance comparisons. Conducting a systematic re-evaluation of representative PIDS under a unified temporal split testing protocol and hyperparameter tuning restricted to the validation set—using publicly available datasets that satisfy auditability, labeling, and calibration requirements—the authors find that most reported performance gains stem from lexical novelty in executable names or paths rather than sophisticated provenance modeling. They propose quantifying dataset semantic signal quality via feature completeness and field entropy, which explain model sensitivity to architectural choices. On three of four widely used datasets, a simple allowlist matches or outperforms learning-based methods; only Theia, exhibiting the strongest semantic signals, effectively reveals model advantages in alert prioritization and node recovery.
📝 Abstract
Provenance-based intrusion detection systems (PIDS) frequently report strong performance, but the conclusions drawn from these results can be highly sensitive to benchmarking choices and evaluation protocols. We investigate this dependency by re-evaluating representative PIDS on public datasets that meet our audit, labeling, and calibration requirements. Focusing primarily on the audited DARPA TC E3 datasets, we apply a unified protocol with temporally separated test periods and validation-only checkpoint and threshold calibration, and ask which architectural claims are empirically supported. We find that alerting success and investigation utility can diverge sharply, as several systems surface attacks without providing enough process-level context to support forensic investigation. On three of the four primary datasets, a simple allowlist built from training executable names and paths matches or exceeds the selected learned baselines on key operating-point metrics, suggesting that much of their measured performance reflects lexical novelty rather than richer provenance modeling. To explain why only some datasets expose architectural differences, we measure semantic signal quality through feature completeness and field entropy. This analysis helps explain why several audited E3 datasets can expose alerting behavior without reliably separating model architectures, while Theia pairs the strongest semantic signal quality with the clearest improvements in ranking and node-level recovery by our reference model. These results show that architectural claims in PIDS should be interpreted together with the benchmark properties and evaluation protocol that produced them.
Problem

Research questions and friction points this paper is trying to address.

provenance-based intrusion detection
benchmarking
evaluation protocol
forensic investigation
semantic signal quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

provenance-based intrusion detection
evaluation protocol
benchmark sensitivity
semantic signal quality
temporal test separation
🔎 Similar Papers
No similar papers found.