🤖 AI Summary
This study addresses the challenge of additively evaluating compound condition effects, such as noise and channel distortion, in speaker verification. We propose a joint error versus marginal error comparison paradigm and construct a quaternary contrastive framework. Through multi-corpus benchmarking and ten-million-scale trial matching analysis, this work systematically reveals misconceptions regarding saturation mechanisms under near-random equal error rate (EER) conditions and the instability of demographic disparity ratios. The project yields a benchmark dataset comprising 4,068 records, quantifies average EER discrepancies across encoders, and provides a fully reproducible toolchain with clearly defined experimental boundaries. These contributions establish a rigorous empirical foundation for assessing model robustness under compound conditions in speaker verification systems.
📝 Abstract
Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does not establish whether their effects add. InterBias-SV organises this question around a four-term comparison: joint error, two marginal errors, and a common reference. Its results artefact contains 4,068 scored records across 17 experiments, 12 encoder labels, and six speech corpora, totalling 12 million trial evaluations. Three experiment families contain the same-corpus terms needed to compute additive contrasts. For labels assigned to speaker-trained encoders, their mean contrasts are +0.0026, +0.0088, and +0.0024 in equal error rate (EER), with larger variation across settings. These descriptive averages do not establish equivalence to additivity: trial matching, checkpoint identity, and parts of the condition metadata remain unverified. We also examine two interpretation problems. Near-chance EER can make additive predictions difficult to interpret, but chance performance is not a hard EER ceiling, and correlation with the prediction does not identify a saturation mechanism. Ratios of demographic gaps are unstable when their clean reference is near zero; absolute gaps provide a more direct summary. The benchmark provides condition definitions, analysis scripts, and explicit requirements for interpretable compound-condition comparisons, while separating recomputable summaries from claims that require further experimental validation.