🤖 AI Summary
This study addresses the lack of interpretability in speaker verification systems by proposing CoLMbo-SV, a model that integrates pretrained acoustic encoders with large language models. Through audio-language modeling and acoustic measurement injection techniques, it generates verifiable analytical reports grounded in acoustic evidence. The authors construct VoxReason, a dedicated dataset, and introduce an evaluation framework that disentangles encoding, influence, and explanation. Experimental results demonstrate that the proposed method achieves an equal error rate (EER) of 0.99% on VoxCeleb1-O, reducing errors by 80%, alongside a numerical grounding score of 0.82. These findings indicate that CoLMbo-SV successfully unifies high verification accuracy with strong interpretability, offering a principled approach to explainable speaker verification.
📝 Abstract
Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-SV}, a speaker language model that combines strong speaker discrimination with structured, acoustically grounded comparison reports. By connecting a pretrained speaker encoder to a language model and supplying explicit acoustic measurements, CoLMbo-SV makes voice comparisons inspectable without restricting verification to the evidence verbalized in its reports. We additionally introduce \textbf{VoxReason}, paired recordings with measured acoustic properties and comparison reports filtered through numerical and qualitative checks, providing supervision for this combined capability. We also develop an evaluation framework that separates what acoustic information a speaker representation encodes, what influences the verification score, and what the generated report discusses. On VoxCeleb1-O, CoLMbo-SV achieves 0.99\% EER, reducing verification error by approximately 80\% relative to the strongest audio-language baseline fine-tuned on VoxReason, while attaining a numerical-grounding score of 0.82. Our analysis further demonstrates that acoustic correctness and decision relevance are distinct properties of an explanation, exposing a gap that numerical-grounding metrics miss. Together, these contributions substantially advance audio-language speaker verification, bring its accuracy toward that of dedicated speaker encoders while adding checkable acoustic reporting, and establish an empirical framework for connecting natural-language explanations to the decisions they explain.