🤖 AI Summary
This study investigates the unclear mechanisms by which auxiliary objectives influence the embedding space in voice anti-spoofing, where they remain vulnerable to generator variations. Building upon the AASIST3 architecture and multi-corpus experiments, this work systematically compares various auxiliary objectives against cross-entropy through embedding space geometric analysis and Equal Error Rate (EER) evaluation. The analysis reveals that cosine consistency gains essentially stem from spatial contraction, while uncovering phenomena of decision-axis dominance and output collapse. Experimental results demonstrate that training without auxiliary objectives consistently outperforms cross-entropy across different architectures. These findings rectify prevalent misconceptions in existing evaluation practices and offer a novel perspective for the design of robust anti-spoofing systems.
📝 Abstract
Speech anti-spoofing countermeasures degrade when the generator, codec or channel changes, and a common remedy is an auxiliary objective that shapes the embedding space; whether it does is invisible to EER, a pure ranking metric. We compare seven such objectives with cross-entropy over 113 runs on five corpora, AASIST3 at three seeds plus four pre-trained detectors, and measure the embedding space of the 24 AASIST3 runs directly. Raw augmentation displacement makes cosine consistency look effective, but the gain is a smaller space, not a more stable one: normalised by the spread, no configuration consistently improves on cross-entropy. Every trained space is dominated by the single decision axis expected for two classes, whose training-set structure does not transfer, and four runs collapse to a near-constant output that displacement rewards and EER reports as poor accuracy. No auxiliary objective keeps an advantage over cross-entropy across architectures, corpora and seeds.