one-class deepfake detection

Design and implement detection systems that identify synthetic or manipulated media by training models only on bona fide (real) examples, typically using density-estimation or likelihood-based scoring to assign anomaly scores to new samples. Build evaluation procedures and decision rules that separate real from fake inputs and emphasize generalisation so the detector recognises forgeries produced by unseen generative methods.

one-classdeepfakedetection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing deepfake detection methods rely heavily on artifacts left by generative models, leading to a significant drop in generalization when confronted with emerging generative architectures and interactive deception scenarios—such as video or voice impersonation—where the core threat lies in deceptive behavior rather than signal-level anomalies. This work breaks from conventional signal-centric paradigms by systematically integrating Speech Act Theory, Grice’s Cooperative Principle, and Cialdini’s Principles of Influence to construct a novel three-tiered analytical framework encompassing speech acts, dialogic interaction, and audience response. By introducing foundational social theories into media forensics, this framework not only exposes the “generalization illusion” inherent in current approaches but also establishes a new pathway for detecting deception in interactive deepfake contexts, while highlighting critical open challenges in the field.

deceptiondeepfake detectiongenerative models

Existing methods for detecting AI-generated images often lack interpretability and rely on implicit assumptions about synthetic artifacts, limiting their robustness under distributional shifts. This work proposes a training-free detection framework that leverages only the statistical properties of authentic images. By integrating multiple untrained statistical descriptors and applying p-value computation combined with classical statistical ensembling techniques—such as Fisher’s method—it constructs an interpretable probabilistic scoring system to assess the consistency of a given image with the distribution of real data. To our knowledge, this is the first approach to establish a universal detection mechanism grounded entirely in the statistics of genuine images, demonstrating strong robustness and flexibility across diverse generative models and cross-domain scenarios.

distributional shiftfake image detectiongenerative models

A Bias-Free Training Paradigm for More General AI-generated Image Detection

Dec 23, 2024
FG
Fabrizio Guillaro
🏛️ University Federico II | Google DeepMind

Existing AI-generated image detectors achieve strong performance on supervised benchmarks but suffer from poor generalization, primarily due to spurious correlations—such as content and formatting biases—in training data, which hinder learning of generator-specific artifacts. To address this, we propose B-Free, a bias-free training paradigm that conditions on real images and employs Stable Diffusion’s reverse sampling to synthesize semantically aligned fake counterparts, thereby establishing the first paired real-fake training framework that decouples artifact learning from content bias. Our method integrates conditional inversion, content-preserving augmentation, and bias-free contrastive learning. Evaluated across 27 diverse generative models—including FLUX and SD3.5—B-Free achieves state-of-the-art generalization, out-of-distribution robustness, and prediction calibration. Code and dataset are publicly released.

Addressing spurious correlations in training data for better performanceDetecting AI-generated images without data bias interferenceImproving generalization and robustness in forensic detectors

Fake It Until You Break It: On the Adversarial Robustness of AI-generated Image Detectors

Oct 02, 2024
SM
Sina Mavali
🏛️ CISPA Helmholtz Center for Information Security | Ruhr University Bochum | University of Tübingen

This study addresses the insufficient adversarial robustness of AI-generated image detectors in real-world settings, where they are vulnerable to black-box attacks and common social media degradations (e.g., JPEG compression, resizing, color distortion), enabling malicious misuse for disinformation and undermining democratic trust. We conduct the first systematic evaluation of mainstream detectors under combined black-box adversarial perturbations and realistic degradations, revealing that state-of-the-art models suffer over 40% accuracy degradation without model access. To mitigate this, we propose a lightweight CLIP-enhanced defense grounded in zero-shot detection and black-box transfer attack modeling—requiring no retraining or fine-tuning. Our method preserves original detection performance while reducing adversarial success rates by 76%, substantially restoring practical utility and robustness on real platforms. This work delivers a deployable, trustworthy solution for AI-generated content authentication.

AI-generated image detectors lack robustness against adversarial attacksCommercial GenAI detection tools are vulnerable to black-box attacksCurrent classifiers fail under real-world conditions and image degradation

Existing AI-generated image detection methods—particularly those based on GANs and diffusion models—exhibit limited generalization to unseen generative models. Method: This paper proposes a language-guided contrastive learning framework that introduces, for the first time, a joint language–vision contrastive supervision mechanism. By leveraging textual labels to enhance visual feature learning, the approach enables zero-shot detection of images from unknown generative models. Specifically, it freezes the CLIP visual encoder and incorporates a learnable text projection head to align multimodal features, thereby improving cross-model generalization. Contribution/Results: Evaluated on four benchmark datasets, the method consistently outperforms all state-of-the-art approaches, achieving an average 12.6% improvement in detection accuracy for unseen generative models. The source code is publicly available.

Address authenticity concerns from synthetic image misuseDetect AI-generated images with improved generalizationEnhance forensic algorithms for diverse generation models

Latest Papers

What's happening recently
View more

This study addresses the high false positive rates and lack of rigorous validation mechanisms in current inductive AI models for detecting synthetic media in forensic contexts. To overcome these limitations, the work introduces abductive reasoning into judicial-grade synthetic media detection for the first time, constructing a fact matrix that synergistically integrates multiple probabilistic detection models with state-of-the-art watermarking techniques such as OpenAI’s SynthID. This approach enables mutual corroboration among diverse detection outputs, significantly reducing false positives while maintaining high true positive recall. The paper also presents the first empirical evaluation of SynthID, revealing complementary strengths among different detection modalities and offering a reliable, interpretable technical pathway suitable for forensic applications.

abductive reasoningfalse positivesforensic AI

AI-generated content poses significant risks in misinformation dissemination and privacy violations; however, existing detection models suffer from poor generalizability, weak cross-model and cross-modal robustness, and limited efficacy against highly manipulated content. Method: This study systematically reviews 24 state-of-the-art works, identifying common challenges in synthetic media detection for the first time, and proposes a unified technical framework centered on multimodal deep learning—integrating CNNs, Vision Transformers (ViTs), and other architectures to enable scalable multimodal fusion. Contribution/Results: We rigorously characterize the failure boundaries of current approaches, establish a theoretical analysis paradigm for generalized detection, and deliver a reproducible, methodology-driven pipeline for robust, cross-domain, and multi-source synthetic content identification.

Current approaches are ineffective for multimodal and highly modified contentDetecting synthetic media struggles with generalization across unseen dataDeveloping robust detection methods against harmful AI-generated media misuse

Existing methods for detecting images generated by diffusion models rely on time-consuming reconstruction and exhibit poor generalization. This work proposes FIND, a novel approach that, for the first time, trains a lightweight binary classifier by perturbing real images with Gaussian noise and labeling them as synthetic samples. FIND directly captures the intrinsic distributional discrepancy between real and synthetic images in terms of their difficulty in Gaussian fitting, without requiring image reconstruction or priors specific to any generative model. The method establishes an end-to-end efficient detection framework that achieves strong performance on the GenImage benchmark, improving detection accuracy by 11.7% while operating 126 times faster than current state-of-the-art approaches—demonstrating a significant advance in balancing universality, efficiency, and practicality.

diffusion-generated image detectiondistributional differenceimage forgery detection

AI-Generated Image Detection: An Empirical Study and Future Research Directions

Nov 04, 2025
NT
Nusrat Tasnim
🏛️ Korea Aerospace University | University of Michigan-Flint

AI-generated images—particularly deepfakes—pose severe threats to multimedia forensics, misinformation detection, and biometric authentication, exacerbating fraud and social engineering risks. Existing detection methods suffer from three key limitations: (i) non-standardized benchmark datasets, (ii) inconsistent training protocols (e.g., end-to-end training, feature freezing, or fine-tuning applied indiscriminately), and (iii) narrow evaluation metrics lacking generalization assessment and interpretability analysis. To address these issues, we propose the first systematic, reproducible benchmark framework for evaluating AI-generated image detection. It uniformly assesses ten state-of-the-art methods across seven diverse datasets spanning GAN- and diffusion-based generators. Our framework introduces multi-dimensional quantitative metrics—including ROC-AUC, class-wise sensitivity, and error rates—alongside interpretability analyses via Grad-CAM and confidence calibration curves. Crucially, we empirically reveal a significant performance gap between in-distribution accuracy and cross-model generalization capability—a previously uncharacterized limitation. This work establishes an empirical foundation and principled methodology for developing robust, interpretable detection systems.

Addressing non-standardized benchmarks for AI-generated image detectionDeveloping robust metrics for generalization and explainability assessmentEvaluating inconsistent training protocols in forensic methods

Hot Scholars

TS

Tejal Shah

Newcastle University
ontologyOWLinformaticsdata integration
XQ

Xueqi Qiu

University of Durham
3D Computer Vision3D ReconstructionDeep Learning
VO

Varun Ojha

Newcastle University
Artificial IntelligenceMachine LearningSignal Processing
XM

Xingyu Miao

Durham University, Department of computer Science
YL

Yang Long

Department of Computer Science, Durham University
Computer VisionMachine LearningArtificial Intelligence