🤖 AI Summary
This study addresses the limited generalization of existing AI-generated image detectors caused by their reliance on static representations by proposing the RED framework. This method leverages multi-scale VQ-VAE reconstruction trajectories alongside the CLIP feature space and, for the first time, exploits a Visual Autoregressive (VAR) model to capture inter-scale token predictability, thereby dynamically aggregating intermediate-stage forensic cues for cross-scale evidence fusion. The architecture employs frozen encoders combined with an adaptive stage-wise weight learning strategy. Evaluated across six benchmarks, RED achieves an average accuracy of 92.5% and a precision of 97.5%. Furthermore, it demonstrates strong robustness against common degradations, significantly outperforming state-of-the-art methods.
📝 Abstract
The rapid evolution of image generators calls for forensic cues that generalize beyond known generation mechanisms. Existing detectors often rely on static image representations or endpoint reconstruction discrepancies, leaving the evolution of intermediate reconstruction stages underexplored. We observe that the relative token predictability of real and generated images can reverse across reconstruction scales, suggesting that intermediate stages may expose forensic evidence overlooked by endpoint comparisons. Motivated by this observation, we propose RED (Reconstruction Evolution Dynamics), a framework that captures transferable forensic cues from coarse-to-fine reconstruction evolution. To our knowledge, RED is the first framework to use scale-wise token predictability to guide forensic evidence aggregation across intermediate reconstruction states. It represents the reconstruction trajectory produced by a frozen multiscale VQ-VAE in the shared feature space of a frozen CLIP encoder. To connect the observed predictability variations with visual evidence, RED learns image-adaptive stage weights from scale-wise token negative log-likelihoods provided by a frozen VAR model. A cross-stage evidence aggregation module then jointly models the original-image representation and the weighted reconstruction features, capturing complementary forensic cues through interactions along the reconstruction trajectory. Experiments on six diverse benchmarks demonstrate that RED achieves the highest average accuracy of 92.5\% and average precision of 97.5\% among the evaluated methods. Further evaluations show strong robustness to common image degradations, supporting the value of reconstruction evolution for generalizable AI-generated image detection. The code will be made publicly available upon acceptance of this paper.