Machine learning for journal entry testing: A type-aware evaluation of anomaly detectors under a review budget

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of traditional anomaly detection evaluation, which often overlooks audit budgets and type distributions, allowing high hit rates to mask critical detection blind spots. We propose a budget-aware evaluation framework based on Fair-Share Type Recall (FSR) and systematically benchmark unsupervised and semi-supervised methods—including PCA, autoencoders, kNN, HBOS, and DeepSAD—on real-world ledger data. Our results demonstrate that the FSR metric significantly alters detector rankings by exposing illusory efficiency driven by the dominance of a single anomaly type. Furthermore, we show that feedback mechanisms may reinforce existing patterns rather than broaden detection coverage. By redefining optimal detector rankings under budget constraints, this work cautions against relying solely on aggregate hit rates and reveals the systemic blind spots they conceal.
📝 Abstract
Journal entry anomaly detectors are commonly evaluated on the full population with ROC-AUC, precision and recall, ignoring the review budget and which anomaly types are found. We propose a type-aware evaluation combining per-type recall, fair-share type recall (FSR), which caps each type's credit at its budget share, type coverage and first-hit rank. We evaluate nine unsupervised detectors, a supervised reference and feedback-driven Deep Semi-Supervised Anomaly Detection (DeepSAD) on four real client ledgers with injected typed anomalies and a public synthetic ledger. On the largest client ledger, principal component analysis (PCA), an autoencoder (AE) and a variational autoencoder (VAE) each place on average 98 anomalies among the first 100 postings, but at least 95.8 belong to one type. FSR instead favours a nearest-neighbour (kNN) detector and changes the top-ranked detector on three of four client ledgers. Representation also matters: one-hot encoding exposes unseen accounts, whereas frequency encoding leaves unseen contra accounts largely undetected. On the public ledger, the Histogram-Based Outlier Score (HBOS) and Empirical Cumulative Distribution-Based Outlier Detection (ECOD) reach all eight markings within 1,386 entries, whereas kNN, the hit leader at 1,000 entries, first reaches cross-linked clearing at rank 4,641, and the supervised row-level reference misses this marking within 1,000 entries. There, the adaptive DeepSAD review protocol raises mean hits per 100 reviews from 40.0 to 68.3 but type coverage only from 2.7 to 3.0. These findings show that high hit rates can conceal systematic blind spots and suggest that feedback can reinforce existing detection patterns without broadening anomaly coverage.
Problem

Research questions and friction points this paper is trying to address.

Journal entry anomaly detection
Review budget
Type-aware evaluation
Anomaly coverage
Detection blind spots
Innovation

Methods, ideas, or system contributions that make the work stand out.

Type-aware evaluation
Anomaly detection
Journal entry testing
Review budget
DeepSAD
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jan Gronewald
aSaarland University, Campus, 66123, Saarbrücken, Germany; bGerman Research Center for Artificial Intelligence (DFKI), Institute for Information Systems (IWi), Campus D3.2, 66123, Saarbrücken, Germany
M
Michel Scherer
aSaarland University, Campus, 66123, Saarbrücken, Germany; bGerman Research Center for Artificial Intelligence (DFKI), Institute for Information Systems (IWi), Campus D3.2, 66123, Saarbrücken, Germany
Nijat Mehdiyev
Nijat Mehdiyev
Senior Researcher, German Research Center for Artificial Intelligence
Explainable Artificial IntelligenceProcess MiningIndustry 4.0