🤖 AI Summary
This study addresses the pressing challenge posed by generative AI–synthesized images to the authenticity of digital media by proposing a detection method based on multimodal handcrafted features and ensemble learning. The discriminative efficacy of features—including DCT, HOG, LBP, GLCM, wavelet transforms, and color histograms—is systematically evaluated on the CIFAKE dataset, and multiple feature sets are fused as input to ensemble models such as LightGBM, XGBoost, and CatBoost. Experimental results demonstrate that feature fusion substantially enhances detection performance. Notably, LightGBM with the combined feature set achieves a PR-AUC of 0.9879, ROC-AUC of 0.9878, F1 score of 0.9447, and Brier score of 0.0414, outperforming single-feature approaches while maintaining high accuracy, interpretability, and computational efficiency.
📝 Abstract
The rapid progress of generative models has enabled the creation of highly realistic synthetic images, raising concerns about authenticity and trust in digital media. Detecting such fake content reliably is an urgent challenge. While deep learning approaches dominate current literature, handcrafted features remain attractive for their interpretability, efficiency, and generalizability. In this paper, we conduct a systematic evaluation of handcrafted descriptors, including raw pixels, color histograms, Discrete Cosine Transform (DCT), Histogram of Oriented Gradients (HOG), Local Binary Patterns (LBP), Gray-Level Co-occurrence Matrix (GLCM), and wavelet features, on the CIFAKE dataset of real versus synthetic images. Using 50,000 training and 10,000 test samples, we benchmark seven classifiers ranging from Logistic Regression to advanced gradient-boosted ensembles (LightGBM, XGBoost, CatBoost). Results demonstrate that LightGBM consistently outperforms alternatives, achieving PR-AUC 0.9879, ROC-AUC 0.9878, F1 0.9447, and a Brier score of 0.0414 with mixed features, representing strong gains in calibration and discrimination over simpler descriptors. Across three configurations (baseline, advanced, mixed), performance improves monotonically, confirming that combining diverse handcrafted features yields substantial benefit. These findings highlight the continued relevance of carefully engineered features and ensemble learning for detecting synthetic images, particularly in contexts where interpretability and computational efficiency are critical.