🤖 AI Summary
This study addresses the challenges of financial fraud detection, including data silos, privacy compliance constraints, and minority class scarcity. To overcome these issues, this work proposes CollaFuse, a collaborative diffusion framework that integrates diffusion models with federated learning to generate high-quality synthetic data through cross-organizational collaboration. A core innovation of this approach lies in transcending the limitations of traditional local data fidelity by instead identifying and leveraging transferable cross-organizational structural information. Experimental evaluations across five datasets demonstrate that, compared to existing baselines, the proposed framework consistently enhances downstream fraud detection performance. Ultimately, this research establishes an effective paradigm for privacy-preserving collaborative modeling across institutions, offering a robust solution for augmenting imbalanced financial data under stringent regulatory requirements.
📝 Abstract
Organizations seek analytical value from AI, yet relevant data are often fragmented across organizations and constrained by privacy. This is acute in financial fraud detection, where rare fraud cases and imbalanced local datasets limit decision-relevant analytics. Federated learning enables collaboration without direct data sharing but does not resolve minority-class scarcity. Synthetic data generation can help, yet lightweight methods are interpolation-bound, while generative models require substantial data and computation. Existing collaborative generative approaches often rely on federated learning, imposing considerable organization-side training burdens. In this paper, we examine CollaFuse as a collaborative diffusion-based alternative for fraud detection and evaluate it across five fraud datasets. Compared with classical oversampling, local generative baselines, and centralized diffusion benchmarks, CollaFuse does not achieve the highest local fidelity but improves downstream fraud detection more consistently across most datasets. These findings suggest that synthetic data create analytical value less through local realism than through transferable cross-organizational structure.