🤖 AI Summary
This study addresses the fragmentation of Vietnamese speech resources, which hinders progress in automatic speech recognition, dialect identification, and deepfake detection, by constructing the first large-scale, multi-domain Vietnamese speech corpus. The dataset integrates authentic and synthetic speech, uniquely providing a unified benchmark that includes transcriptions, speaker identities, annotations for five dialects, and natural code-switching. Furthermore, controlled spoofed speech is generated using four open-source and commercial systems to facilitate zero-shot evaluation. Experimental results reveal significant vulnerabilities in existing detectors under cross-generator pairings and high-similarity conditions. By establishing this comprehensive resource and exposing critical limitations in current approaches, this work sets a new standard for intelligent Vietnamese speech processing and security verification.
📝 Abstract
Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos. To our knowledge, it is the first large-scale Vietnamese corpus to jointly provide transcripts, consistent speaker identities, five dialect groups, and naturally occurring Vietnamese--English code-switching, which constitutes nearly half of the corpus by duration. We further create over 3.1K hours of spoof speech with four open-source and commercial synthesis systems. Every spoof is conditioned on a verified speaker reference and paired with a transcript- and speaker-matched bona fide utterance, enabling unique controlled evaluation with reduced lexical and identity confounds. Zero-shot evaluation of five pretrained multilingual detectors reveals striking brittleness: EER greatly varies across detector--generator pairings, while recent multilingual detector DFA-1B degrades from 16.3% to 33.6% as speaker similarity increases. Dialect-stratified results expose further model-dependent disparities. By unifying natural linguistic diversity with controlled spoof generation, VietPrism provides a challenging foundation for Vietnamese speech modeling and trustworthy audio-deepfake detection.