π€ AI Summary
This study addresses the challenges of missing negative samples and underutilized network information in tax fraud detection by proposing a Bayesian positive-unlabeled (PU) learning framework that integrates covariates with multiplex networks. Methodologically, a Gaussian process model is constructed via product-of-experts, pioneering the incorporation of multiplex network structures into a Bayesian PU architecture. Furthermore, theoretical bounds on the false-negative rate are derived, enabling the decoupling of risk-driving factors. Experimental results demonstrate the frameworkβs superior ranking performance. In real-world case studies, the approach independently identified two non-compliant enterprises, facilitated the targeting of four additional entities for review, and quantified the scale of potentially undetected fraud.
π Abstract
Tax fraud remains a central challenge for public revenue authorities worldwide, imposing fiscal losses estimated to reach up to 1 trillion euros annually in the EU alone. Fraudulent firms intentionally manipulate reported figures to conceal their activity. Business networks can provide complementary information and reveal fraud that covariates alone may miss. We propose a Bayesian positive-unlabeled (PU) classification framework that combines firm-level covariates with multilayer network information. As a policy relevant case study, we apply the framework to firms in a regulated segment of the Greek energy market, a sector exposed to excise tax evasion. Confirmed fraud labels exist only for audited cases, while the remaining observations are unlabeled rather than verified compliant. We develop a Bayesian Gaussian process (GP) classifier integrating covariates with multilayer network information through a Product of Experts (PoE) construction and incorporating a nondetection probability for missed positives. We derive identification bounds for the nondetection rate under an anchor condition and show that, for a fixed latent risk function, a common nondetection rate affects calibration but not ranking. Simulations show strong ranking performance relative to PU and network-based alternatives, while illustrating the difficulty of estimating the nondetection rate. In real audit data, the method identifies high risk firms with posterior uncertainty, estimates the number of undetected fraudulent cases, and identifies whether risk is driven by covariates, networks, or both. Of the six highest ranked firms, the two that had already been reinspected independently were both confirmed as noncompliant, providing an independent validation of these cases. The remaining four were selected by the tax authority for follow-up inspection.