Hybrid Data can Enhance the Utility of Synthetic Data for Training Anti-Money Laundering Models

📅 2025-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In anti-money laundering (AML) model training, access to real transaction data is severely constrained by privacy regulations and compliance requirements, while purely synthetic data suffers from limited generalization capability. Method: This paper proposes a hybrid data construction paradigm that integrates differentially private generative synthetic data with publicly available real-world financial network topologies and statistical characteristics. Crucially, it introduces, for the first time, a graph neural network (GNN)-driven multi-source feature alignment mechanism into the synthetic data augmentation pipeline, enabling joint modeling of structural fidelity and semantic consistency. Results: Empirical evaluation demonstrates that AML models trained on the hybrid data significantly outperform those trained on purely synthetic baselines across key metrics—including F1-score and false positive rate—while exhibiting strong practical utility and cross-scenario generalizability. The approach establishes a reproducible, scalable framework for financial risk modeling in privacy-sensitive domains.

Technology Category

Machine Learning: PrivacyNatural Language Processing: GenerationHumans and AI: Human-in-the-loop Machine Learning

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSecurity and Privacy: Data transparency and provenance
📝 Abstract
Money laundering is a critical global issue for financial institutions. Automated Anti-money laundering (AML) models, like Graph Neural Networks (GNN), can be trained to identify illicit transactions in real time. A major issue for developing such models is the lack of access to training data due to privacy and confidentiality concerns. Synthetically generated data that mimics the statistical properties of real data but preserves privacy and confidentiality has been proposed as a solution. However, training AML models on purely synthetic datasets presents its own set of challenges. This article proposes the use of hybrid datasets to augment the utility of synthetic datasets by incorporating publicly available, easily accessible, and real-world features. These additions demonstrate that hybrid datasets not only preserve privacy but also improve model utility, offering a practical pathway for financial institutions to enhance AML systems.
Problem

Research questions and friction points this paper is trying to address.

Lack of real training data for AML models due to privacy concerns
Purely synthetic data presents challenges for training effective AML systems
Hybrid datasets can enhance synthetic data utility while preserving privacy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid datasets augment synthetic data utility
Incorporating publicly available real-world features
Preserving privacy while improving model performance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rachel Chung
College of William and Mary
P
Pratyush Nidhi Sharma
The University of Alabama
Mikko Siponen
Mikko Siponen
Professor of Business Cyber Security, Professor of MIS, the University of Alabama
Cybersecurity Managementcybercrimebehavioral information securityusable securityMIS
R
Rohit Vadodaria
CGI Federal
L
Luke Smith
College of William and Mary