Where the Numbers Come From: Auditing Evaluation in Provenance-Based Intrusion Detection

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses metric distortions in existing intrusion detection evaluations caused by label selection, checkpointing, and encoding deficiencies. We propose the STRICT framework, which systematically isolates and quantifies three categories of measurement effects by auditing nine implementations through code execution, Word2Vec-based configuration analysis, and linear encoder testing. The investigation reveals deep-seated flaws such as fixed-alert comparisons and buffer reuse, while validating alert precision and attribute sufficiency. Following these corrections, detection accuracy changes substantially, thereby delimiting the applicability scope of prior conclusions. Ultimately, this work establishes a new paradigm for constructing rigorous intrusion detection evaluation benchmarks.
📝 Abstract
Reproducing a provenance-based intrusion detector's score does not establish what that score says about its emitted alarms or the information its encoder uses. We audit nine released implementations, execute four detectors using their own code, and isolate three measurement effects. First, a fixed-alert comparison separates label choice from neighbourhood credit: ThreaTrace reports precision 0.938 with neighbourhood credit, although only seven of its 994 alarms carry its own attack label. Second, removing test-label checkpoint selection lowers attack detection precision (ADP) by 0.16 to 0.33 across four forty-member word2vec configurations without changing detector order. Third, a buffer-reuse defect gives a linear encoder unintended degree-dependent inputs. Correcting it lowers type-only ADP in every identical-input initialization pair on two hosts, while historical word2vec effects depend on the host. These findings qualify claims from the inspected implementations about alarm precision, performance magnitude and static-attribute sufficiency. STRICT connects them to six checkable reporting requirements. Because the comparisons condition on benchmark targets, they neither validate those labels nor establish a universal detector ranking.
Problem

Research questions and friction points this paper is trying to address.

provenance-based intrusion detection
evaluation auditing
measurement effects
alarm precision
benchmark reproducibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Provenance-based Intrusion Detection
Evaluation Auditing
Measurement Effects
Neighbourhood Credit
STRICT Framework
🔎 Similar Papers
No similar papers found.