PC-Mix: Partial-Component Audio Spoofing Detection under Mixed Speech and Environmental Sound Conditions

📅 2026-07-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing research on audio spoofing detection often overlooks real-world scenarios where speech coexists with environmental sounds and may be partially manipulated. To address this gap, this work introduces PC-Mix, the first dataset specifically designed for partial-component spoofing detection in mixed audio, featuring realistic scenarios with locally forged environmental sounds. We further propose a joint learning framework that operates across both speech and environmental sound components. Through a unified evaluation protocol, our experiments demonstrate that spoofing detection under mixed conditions is significantly more challenging, and models trained under target-matched conditions substantially outperform those directly transferred from single-component settings. This study fills a critical research void in partial spoofing detection involving environmental sounds and mixed acoustic conditions.
📝 Abstract
Recent studies on partial audio spoofing mainly focus on studio-recorded speech with temporal localization of spoofed segments. However, these studies often overlook realistic conditions where spoofed and bonafide segments simultaneously coexist across speech and environmental sound components. In this paper, we present PC-Mix, the first dataset for partial-component spoofing detection, where either or both audio components may be partially spoofed. In PC-Mix, bonafide and partially spoofed environmental-sound components are first constructed and mixed with speech signals from an existing partial-spoof dataset, producing audio in which either or both components may be locally manipulated. This design addresses two major gaps in existing partial spoofing benchmarks: the lack of realistic environmental sounds in speech partial spoofing scenarios and the absence of partial spoofing detection for environmental sound components. We further establish standardized evaluation protocols and design a joint learning framework to optimize spoofing detection across speech, environmental sound, and mixed audio. Experiments highlight the increased difficulty introduced by mixed conditions. The results demonstrate that training under matched target conditions is more effective than directly transferring models trained on speech or environmental sound components.
Problem

Research questions and friction points this paper is trying to address.

partial audio spoofing
environmental sound
mixed audio
spoofing detection
realistic conditions
Innovation

Methods, ideas, or system contributions that make the work stand out.

partial-component spoofing
mixed audio
environmental sound
joint learning framework
audio spoofing detection
🔎 Similar Papers
2024-04-22arXiv.orgCitations: 25
Z
Zhenshan Zhang
Digital Innovation Research Center, Duke Kunshan University, Kunshan, China
X
Xueping Zhang
Digital Innovation Research Center, Duke Kunshan University, Kunshan, China
L
Linxi Li
OfSpectrum, Inc., Los Angeles, USA
Y
Yechen Wang
OfSpectrum, Inc., Los Angeles, USA
M
Ming Li
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China