Learning as Deepfakes Evolve: RF-Prompt for Continual Audio Deepfake Detection

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of catastrophic forgetting and the entanglement of domain shifts with mechanism variations in audio deepfake detection when continuously learning new forgery mechanisms. To this end, it proposes the RAMI evaluation protocol to simulate realistic incremental scenarios and introduces RF-Prompt, a method featuring an asymmetric prompt architecture. Specifically, RF-Prompt inherits forgery experts through shared authentic prompts and orthogonal residual expansion, combined with input-adaptive soft fusion to achieve incremental detection without requiring task identifiers during inference. By synergizing continual learning and mixture-of-experts techniques, the proposed approach achieves an average equal error rate (EER) of 10.11% under the RAMI protocol, significantly outperforming baseline methods and demonstrating its effectiveness and robustness.
📝 Abstract
Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech. Existing dataset-incremental evaluation changes both real-speech domains and deepfake mechanisms, making their effects difficult to distinguish. We construct five task organizations over identical training, development, and evaluation pools to study these factors under a controlled sample budget. Our proposed Real-Anchored Mechanism-Incremental (RAMI) protocol reflects the practical setting in which available real speech provides a recurring mixed-domain reference while new deepfake mechanisms arrive incrementally. We further propose RF-Prompt, an asymmetric continual prompt-learning method that preserves reusable real-speech knowledge through a shared real prompt and expands mechanism-specific knowledge through inherited fake experts with orthogonal residuals. Input-adaptive soft fusion combines the accumulated experts into a fixed number of injected tokens without requiring task identity at inference. On RAMI, RF-Prompt achieves 10.110% average EER and 10.370% pooled EER, outperforming all evaluated continual-learning baselines. Across the five controlled protocols, RAMI yields the lowest common-average and pooled EER. Component ablations, limited-data experiments, and cross-backbone evaluations further validate the proposed design.
Problem

Research questions and friction points this paper is trying to address.

Continual learning
Audio deepfake detection
Dataset-incremental evaluation
Mechanism-incremental
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continual Audio Deepfake Detection
Prompt Learning
RAMI Protocol
Input-Adaptive Soft Fusion
Orthogonal Residuals
🔎 Similar Papers
No similar papers found.