Score
Design and implement detection systems that identify synthetic or manipulated media by training models only on bona fide (real) examples, typically using density-estimation or likelihood-based scoring to assign anomaly scores to new samples. Build evaluation procedures and decision rules that separate real from fake inputs and emphasize generalisation so the detector recognises forgeries produced by unseen generative methods.
Existing deepfake detection methods rely heavily on artifacts left by generative models, leading to a significant drop in generalization when confronted with emerging generative architectures and interactive deception scenarios—such as video or voice impersonation—where the core threat lies in deceptive behavior rather than signal-level anomalies. This work breaks from conventional signal-centric paradigms by systematically integrating Speech Act Theory, Grice’s Cooperative Principle, and Cialdini’s Principles of Influence to construct a novel three-tiered analytical framework encompassing speech acts, dialogic interaction, and audience response. By introducing foundational social theories into media forensics, this framework not only exposes the “generalization illusion” inherent in current approaches but also establishes a new pathway for detecting deception in interactive deepfake contexts, while highlighting critical open challenges in the field.
Existing methods for detecting AI-generated images often lack interpretability and rely on implicit assumptions about synthetic artifacts, limiting their robustness under distributional shifts. This work proposes a training-free detection framework that leverages only the statistical properties of authentic images. By integrating multiple untrained statistical descriptors and applying p-value computation combined with classical statistical ensembling techniques—such as Fisher’s method—it constructs an interpretable probabilistic scoring system to assess the consistency of a given image with the distribution of real data. To our knowledge, this is the first approach to establish a universal detection mechanism grounded entirely in the statistics of genuine images, demonstrating strong robustness and flexibility across diverse generative models and cross-domain scenarios.
Existing AI-generated image detectors achieve strong performance on supervised benchmarks but suffer from poor generalization, primarily due to spurious correlations—such as content and formatting biases—in training data, which hinder learning of generator-specific artifacts. To address this, we propose B-Free, a bias-free training paradigm that conditions on real images and employs Stable Diffusion’s reverse sampling to synthesize semantically aligned fake counterparts, thereby establishing the first paired real-fake training framework that decouples artifact learning from content bias. Our method integrates conditional inversion, content-preserving augmentation, and bias-free contrastive learning. Evaluated across 27 diverse generative models—including FLUX and SD3.5—B-Free achieves state-of-the-art generalization, out-of-distribution robustness, and prediction calibration. Code and dataset are publicly released.
This study addresses the insufficient adversarial robustness of AI-generated image detectors in real-world settings, where they are vulnerable to black-box attacks and common social media degradations (e.g., JPEG compression, resizing, color distortion), enabling malicious misuse for disinformation and undermining democratic trust. We conduct the first systematic evaluation of mainstream detectors under combined black-box adversarial perturbations and realistic degradations, revealing that state-of-the-art models suffer over 40% accuracy degradation without model access. To mitigate this, we propose a lightweight CLIP-enhanced defense grounded in zero-shot detection and black-box transfer attack modeling—requiring no retraining or fine-tuning. Our method preserves original detection performance while reducing adversarial success rates by 76%, substantially restoring practical utility and robustness on real platforms. This work delivers a deployable, trustworthy solution for AI-generated content authentication.
Existing AI-generated image detection methods—particularly those based on GANs and diffusion models—exhibit limited generalization to unseen generative models. Method: This paper proposes a language-guided contrastive learning framework that introduces, for the first time, a joint language–vision contrastive supervision mechanism. By leveraging textual labels to enhance visual feature learning, the approach enables zero-shot detection of images from unknown generative models. Specifically, it freezes the CLIP visual encoder and incorporates a learnable text projection head to align multimodal features, thereby improving cross-model generalization. Contribution/Results: Evaluated on four benchmark datasets, the method consistently outperforms all state-of-the-art approaches, achieving an average 12.6% improvement in detection accuracy for unseen generative models. The source code is publicly available.
This study addresses the high false positive rates and lack of rigorous validation mechanisms in current inductive AI models for detecting synthetic media in forensic contexts. To overcome these limitations, the work introduces abductive reasoning into judicial-grade synthetic media detection for the first time, constructing a fact matrix that synergistically integrates multiple probabilistic detection models with state-of-the-art watermarking techniques such as OpenAI’s SynthID. This approach enables mutual corroboration among diverse detection outputs, significantly reducing false positives while maintaining high true positive recall. The paper also presents the first empirical evaluation of SynthID, revealing complementary strengths among different detection modalities and offering a reliable, interpretable technical pathway suitable for forensic applications.
AI-generated content poses significant risks in misinformation dissemination and privacy violations; however, existing detection models suffer from poor generalizability, weak cross-model and cross-modal robustness, and limited efficacy against highly manipulated content. Method: This study systematically reviews 24 state-of-the-art works, identifying common challenges in synthetic media detection for the first time, and proposes a unified technical framework centered on multimodal deep learning—integrating CNNs, Vision Transformers (ViTs), and other architectures to enable scalable multimodal fusion. Contribution/Results: We rigorously characterize the failure boundaries of current approaches, establish a theoretical analysis paradigm for generalized detection, and deliver a reproducible, methodology-driven pipeline for robust, cross-domain, and multi-source synthetic content identification.
Existing methods for detecting images generated by diffusion models rely on time-consuming reconstruction and exhibit poor generalization. This work proposes FIND, a novel approach that, for the first time, trains a lightweight binary classifier by perturbing real images with Gaussian noise and labeling them as synthetic samples. FIND directly captures the intrinsic distributional discrepancy between real and synthetic images in terms of their difficulty in Gaussian fitting, without requiring image reconstruction or priors specific to any generative model. The method establishes an end-to-end efficient detection framework that achieves strong performance on the GenImage benchmark, improving detection accuracy by 11.7% while operating 126 times faster than current state-of-the-art approaches—demonstrating a significant advance in balancing universality, efficiency, and practicality.
AI-generated images—particularly deepfakes—pose severe threats to multimedia forensics, misinformation detection, and biometric authentication, exacerbating fraud and social engineering risks. Existing detection methods suffer from three key limitations: (i) non-standardized benchmark datasets, (ii) inconsistent training protocols (e.g., end-to-end training, feature freezing, or fine-tuning applied indiscriminately), and (iii) narrow evaluation metrics lacking generalization assessment and interpretability analysis. To address these issues, we propose the first systematic, reproducible benchmark framework for evaluating AI-generated image detection. It uniformly assesses ten state-of-the-art methods across seven diverse datasets spanning GAN- and diffusion-based generators. Our framework introduces multi-dimensional quantitative metrics—including ROC-AUC, class-wise sensitivity, and error rates—alongside interpretability analyses via Grad-CAM and confidence calibration curves. Crucially, we empirically reveal a significant performance gap between in-distribution accuracy and cross-model generalization capability—a previously uncharacterized limitation. This work establishes an empirical foundation and principled methodology for developing robust, interpretable detection systems.