🤖 AI Summary
This study addresses the critical issue that test-set model selection rules exert a decisive influence on rankings in graph anomaly detection, potentially introducing evaluation bias. Through a preregistered auditing experiment across eight graphs, this work compares the performance of DEMO, NSReg, and OUTPOST under two selection criteria—“best epoch” versus “validation set”—while incorporating deployable validation strategies and multiple statistical tests. The findings reveal scoring biases induced by benchmark construction practices, demonstrating that altering selection rules can directly change the top-ranked model. Notably, such bias is pronounced on small datasets but less evident in real-world fraud scenarios. Accordingly, this paper proposes a standardized reporting checklist to provide normative guidance for evaluating open-set graph anomaly detection methods.
📝 Abstract
Open-set graph anomaly detection trains on a few labeled anomalies from one class and must also find anomaly classes that were never labeled. Published results share three conventions: the test score is read at the best epoch on the test set, baseline numbers are copied from earlier papers, and most anomalies are minority classes relabeled as anomalous. We ask how much of the reported ranking these conventions decide. We re-run two recent methods, DEMO and NSReg, together with OUTPOST, a small first-order detector built for this study. All three use one protocol with identical seeds and splits on eight graphs (seven for the baselines, which cannot run on ogbn-mag), ten seeds each, and every run is scored under both the best-epoch rule and a deployable validation rule. Before the runs that test them, we registered 40 predictions. Three findings hold. First, the rule changes the leader: under the best-epoch rule, OUTPOST and NSReg each lead three of seven graphs, while under the validation rule, NSReg leads five. Second, the best-epoch bonus depends on how the benchmark was built: 0.045--0.080 AUC-ROC on the three small relabeled-class graphs and 0.002--0.014 on the three real fraud graphs. Third, pseudo-labeling in OUTPOST is worth 0.038--0.065 AUC-ROC on the same three graphs but gives no benefit on any real fraud graph. We also show that a 0.002 tie band for hyperparameter selection lies below the paired standard error on all six graphs tested, even at ten seeds. Twelve of our 40 predictions were falsified, and we report them. We close with a short reporting checklist.