🤖 AI Summary
To address the limitations of existing approaches in credit card fraud detection—namely, insufficient accuracy, poor interpretability, and inadequate real-time performance—this paper proposes an end-to-end unsupervised learning framework. Methodologically, it integrates multi-source financial data to construct temporal behavioral features; introduces a novel composite risk scoring mechanism that fuses anomaly scores from Isolation Forest, One-Class SVM, and a deep autoencoder, augmented with burst- and frequency-based spending indicators; and employs PCA-based visualization coupled with DBSCAN density clustering to enable interpretable localization of fraudulent transactions. Evaluated on real-world transaction data, the framework achieves precise identification of 1–2% high-risk transactions, significantly improving detection priority for high-risk cardholders and merchants. It supports real-time anti-fraud decision-making while maintaining both high detection efficacy and operational interpretability.
📝 Abstract
The rise of digital payments has accelerated the need for intelligent and scalable systems to detect fraud. This research presents an end-to-end, feature-rich machine learning framework for detecting credit card transaction anomalies and fraud using real-world data. The study begins by merging transactional, cardholder, merchant, and merchant category datasets from a relational database to create a unified analytical view. Through the feature engineering process, we extract behavioural signals such as average spending, deviation from historical patterns, transaction timing irregularities, and category frequency metrics. These features are enriched with temporal markers such as hour, day of week, and weekend indicators to expose all latent patterns that indicate fraudulent behaviours. Exploratory data analysis (EDA) reveals contextual transaction trends across all the dataset features. Using the transactional data, we train and evaluate a range of unsupervised models: Isolation Forest, One Class SVM, and a deep autoencoder trained to reconstruct normal behavior. These models flag the top 1% of reconstruction errors as outliers. PCA visualizations illustrate each model’s ability to separate anomalies into a two-dimensional latent space. We further segment the transaction landscape using K-Means clustering and DBSCAN to identify dense clusters of normal activity and isolate sparse, suspicious regions. Finally, we propose a composite risk score by aggregating binary flags from all anomaly detectors, unexpected spend indicators, rapid‐use events, and high‐frequency “spending sprees”. This score highlights the riskiest cardholders and merchants, enabling prioritized investigation. Our framework detects approximately 1–2% of transactions as anomalies and effectively surfaces high‐risk entities, demonstrating the power of unsupervised analytics for real-time fraud surveillance in dynamic financial ecosystems.