🤖 AI Summary
This study addresses the limitations of traditional anomaly detection evaluation, which often overlooks audit budgets and type distributions, allowing high hit rates to mask critical detection blind spots. We propose a budget-aware evaluation framework based on Fair-Share Type Recall (FSR) and systematically benchmark unsupervised and semi-supervised methods—including PCA, autoencoders, kNN, HBOS, and DeepSAD—on real-world ledger data. Our results demonstrate that the FSR metric significantly alters detector rankings by exposing illusory efficiency driven by the dominance of a single anomaly type. Furthermore, we show that feedback mechanisms may reinforce existing patterns rather than broaden detection coverage. By redefining optimal detector rankings under budget constraints, this work cautions against relying solely on aggregate hit rates and reveals the systemic blind spots they conceal.
📝 Abstract
Journal entry anomaly detectors are commonly evaluated on the full population with ROC-AUC, precision and recall, ignoring the review budget and which anomaly types are found. We propose a type-aware evaluation combining per-type recall, fair-share type recall (FSR), which caps each type's credit at its budget share, type coverage and first-hit rank. We evaluate nine unsupervised detectors, a supervised reference and feedback-driven Deep Semi-Supervised Anomaly Detection (DeepSAD) on four real client ledgers with injected typed anomalies and a public synthetic ledger. On the largest client ledger, principal component analysis (PCA), an autoencoder (AE) and a variational autoencoder (VAE) each place on average 98 anomalies among the first 100 postings, but at least 95.8 belong to one type. FSR instead favours a nearest-neighbour (kNN) detector and changes the top-ranked detector on three of four client ledgers. Representation also matters: one-hot encoding exposes unseen accounts, whereas frequency encoding leaves unseen contra accounts largely undetected. On the public ledger, the Histogram-Based Outlier Score (HBOS) and Empirical Cumulative Distribution-Based Outlier Detection (ECOD) reach all eight markings within 1,386 entries, whereas kNN, the hit leader at 1,000 entries, first reaches cross-linked clearing at rank 4,641, and the supervised row-level reference misses this marking within 1,000 entries. There, the adaptive DeepSAD review protocol raises mean hits per 100 reviews from 40.0 to 68.3 but type coverage only from 2.7 to 3.0. These findings show that high hit rates can conceal systematic blind spots and suggest that feedback can reinforce existing detection patterns without broadening anomaly coverage.