Temporal Anchors and Editing Sensitivity in Partial Speech Spoofing: A Controlled Study

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the false alarms and missed detections in partial speech forgery detection caused by confusion between benign edits and synthetic content. We propose a diagnostic framework based on frozen WavLM features. Methodologically, temporal anchors and boundary consistency training are introduced to analyze detector sensitivity to edits, while non-linguistic unit control and source-locking evaluation mechanisms enable forgery localization at frame, phoneme, and word granularities. Experiments reveal that phoneme-level error rates exceed those at the soft-frame level. Furthermore, data augmentation significantly reduces false alarms on authentic splices but may increase missed detections of synthetic segments, whereas consistency training yields no universal gains. This work systematically uncovers the differential impacts of augmentation strategies on detection performance.
📝 Abstract
Partial-spoof detectors must reject synthetic content while accepting benign edits. We diagnose temporal anchors and boundary-consistency training using frozen WavLM features, nonlinguistic unit controls, and source-locked evaluation. Genuine-genuine and genuine-fake splices contrast editing false alarms with synthetic-content misses; they do not isolate a unique causal artifact. With layer-6 features, phone units have higher PartialSpoof evaluation frame EER than soft frames (11.18% versus 9.59%). In 500 retrospective PS-eval cases, augmentation reduces genuine-splice false alarms; synthetic-core misses increase at 5% PS-development FPR but not demonstrably at 1%. A 200-case Llama A follow-up retains the false-alarm reduction but does not establish a miss-rate increase. Constructed-case ranking can improve while fixed-threshold misses rise. Consistency gives no uniform gain, and encoder-layer rankings change across corpora. We report frame and event localization with three seeds, including word units, and treat PS evaluation as retrospective. Audit code, selected training scripts, and result summaries are available at https://github.com/mysxs/partial-spoof-diagnostics.
Problem

Research questions and friction points this paper is trying to address.

partial speech spoofing
spoof detection
false alarm
miss rate
speech editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Partial Speech Spoofing
Temporal Anchors
Boundary-consistency Training
WavLM Features
Nonlinguistic Unit Controls
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiaosu Su
Institute of Information Engineering, Chinese Academy of Sciences, Beijing 100085, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing 100085, China
Yun Cao
Yun Cao
researcher, tencent
CVGANs
Y
Yiping Ni
Institute of Information Engineering, Chinese Academy of Sciences, Beijing 100085, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing 100085, China
X
Xiaowei Yi
Institute of Information Engineering, Chinese Academy of Sciences, Beijing 100085, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing 100085, China