🤖 AI Summary
This study addresses the predominant focus on English in existing multilingual speech anonymization research and the lack of systematic evaluation regarding attacker automatic speaker verification (ASV) behavior. We construct a multilingual voice conversion dataset to systematically investigate, for the first time, the attack mechanisms of acoustic- and content-oriented attackers in cross-lingual scenarios, while analyzing how linguistic utility influences their effectiveness. Experimental results demonstrate that acoustic-oriented attackers generally outperform content-oriented ones; however, this performance gap narrows significantly when complete linguistic information is available. Furthermore, the proposed dataset effectively enhances attack performance and mitigates cross-lingual discrepancies. This work establishes a novel paradigm for the security evaluation of anonymized speech that jointly considers privacy protection and utility preservation.
📝 Abstract
Attacker ASV systems for voice anonymization have been studied primarily in English, leaving their behavior in multilingual settings largely unexplored. Conventional ASV has shown that both acoustic and contextual information are important for multilingual speaker verification. Inspired by this, we investigate whether the same holds for attacker ASV on anonymized speech. We evaluate both acoustic- and content-oriented attackers on multilingual anonymized speech and construct a multilingual voice-converted dataset to improve cross-lingual generalization. Our results show that attacker effectiveness depends on the linguistic utility of the anonymized speech. Overall, acoustic-oriented attackers achieve better performance. However, when linguistic information is well preserved, the performance gap between content- and acoustic-oriented attackers narrows compared with conditions involving stronger speech distortion. The multilingual voice-converted dataset further improves performance and partially reduces the cross-lingual gap. These findings highlight the need for more comprehensive attacker modeling and evaluation protocols that consider both privacy and utility, rather than relying on a attacker strategy\footnote{Full code and pretrained models and MultiVC Dataset link are available at: https://github.com/monkeyDarefeen/DAST