CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of interpretability in speaker verification systems by proposing CoLMbo-SV, a model that integrates pretrained acoustic encoders with large language models. Through audio-language modeling and acoustic measurement injection techniques, it generates verifiable analytical reports grounded in acoustic evidence. The authors construct VoxReason, a dedicated dataset, and introduce an evaluation framework that disentangles encoding, influence, and explanation. Experimental results demonstrate that the proposed method achieves an equal error rate (EER) of 0.99% on VoxCeleb1-O, reducing errors by 80%, alongside a numerical grounding score of 0.82. These findings indicate that CoLMbo-SV successfully unifies high verification accuracy with strong interpretability, offering a principled approach to explainable speaker verification.
📝 Abstract
Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-SV}, a speaker language model that combines strong speaker discrimination with structured, acoustically grounded comparison reports. By connecting a pretrained speaker encoder to a language model and supplying explicit acoustic measurements, CoLMbo-SV makes voice comparisons inspectable without restricting verification to the evidence verbalized in its reports. We additionally introduce \textbf{VoxReason}, paired recordings with measured acoustic properties and comparison reports filtered through numerical and qualitative checks, providing supervision for this combined capability. We also develop an evaluation framework that separates what acoustic information a speaker representation encodes, what influences the verification score, and what the generated report discusses. On VoxCeleb1-O, CoLMbo-SV achieves 0.99\% EER, reducing verification error by approximately 80\% relative to the strongest audio-language baseline fine-tuned on VoxReason, while attaining a numerical-grounding score of 0.82. Our analysis further demonstrates that acoustic correctness and decision relevance are distinct properties of an explanation, exposing a gap that numerical-grounding metrics miss. Together, these contributions substantially advance audio-language speaker verification, bring its accuracy toward that of dedicated speaker encoders while adding checkable acoustic reporting, and establish an empirical framework for connecting natural-language explanations to the decisions they explain.
Problem

Research questions and friction points this paper is trying to address.

Speaker Verification
Explainability
Acoustic Grounding
Audio-Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speaker Verification
Explainable AI
Audio-Language Model
Acoustic Grounding
Evaluation Framework
🔎 Similar Papers
No similar papers found.