Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出边界和段内学习(BISL)方法,通过建模相邻帧特征差异及连续真实与伪造段落的总体特性,有效解决部分音频深度伪造定位难题。
📝 Abstract
Partial audio deepfakes manipulate only selected speech regions, making them difficult to be localized. Existing methods exploit boundary cues for partial deepfake localization, but primarily focus on identifying boundary positions rather than modeling the feature changes that characterize authenticity transitions. Meanwhile, the internal characteristics of continuous bona fide and spoofed segments remain underexplored. In this paper, we propose Boundary and Intra-Segment Learning (BISL), which introduces boundary learning to model feature differences between adjacent frames and distinguish authenticity transitions from general acoustic variations. In addition, intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments while enhancing feature consistency within each segment. By jointly learning frame, boundary, and segment information, BISL enables more effective fine-grained partial audio deepfake localization. Experiments on multiple localization benchmarks show that BISL achieves an EER of 2.52\% and an F1-score of 97.40\% on PartialSpoof, outperforming the compared methods, while maintaining competitive performance on HAD and improved cross-dataset performance on LPS. The code will be made publicly available upon acceptance.
Problem

Research questions and friction points this paper is trying to address.

partial audio deepfakes
localization
boundary cues
feature changes
intra-segment characteristics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Boundary and Intra-Segment Learning
partial audio deepfake localization
feature differences
authenticity transitions
segment characteristics
🔎 Similar Papers
No similar papers found.
Z
Zhe Ye
Guangdong Key Lab of Information Security, School of Computer Science and Engineering, Sun Yat-sen University, China; Nanyang Technological University, Singapore
Xiangui Kang
Xiangui Kang
Professor Xiangui Kang, Sun Yat-Sen University, China
multimedia signal processingcommunication and game theory
M
Minhua Huang
China Mobile Internet Corporation, China
K
Kai Wu
China Mobile Internet Corporation, China
Kong Aik Lee
Kong Aik Lee
The Hong Kong Polytechnic University, Hong Kong
Speaker and Spoken Language RecognitionSpeech ProcessingDigital Signal ProcessingSubband
C
Chng Eng Siong
Nanyang Technological University, Singapore