TS-SP: Learning Speaker-Preserving Representations in Audio Large Language Models

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that existing large audio models struggle to effectively leverage speaker identity information for verification. To this end, it proposes a two-stage LoRA fine-tuning framework based on Qwen2.5-Omni-7B that decouples acoustic encoding from semantic comparison while keeping pretrained weights frozen. The first stage enhances the representational capacity of the audio encoder, whereas the second stage trains the language model for speaker matching via supervised learning. This strategy endows large models with speaker verification capabilities in a parameter-efficient manner. Experimental results demonstrate that the equal error rate (EER) on the Vox1-O dataset decreases from 7.01% to 4.37% compared to the baseline, and the approach maintains robust generalization performance even under unseen prompts.
📝 Abstract
Audio large language models (ALLMs) can understand speech content, yet their ability to use speaker identity for verification remains limited. We propose TS-SP (Two-Stage Speaker Preservation), a parameter-efficient framework for learning speaker-preserving representations and making them accessible to an ALLM's language-model component. We instantiate and evaluate TS-SP on Qwen2.5-Omni-7B. First, we adapt the audio encoder with speaker identity supervision. We then freeze the adapted encoder and train the language model to compare speakers. Both stages use low-rank adaptation (LoRA), keeping the pretrained base weights fixed. On Vox1-O, TS-SP reduces the equal error rate (EER) from 7.01\% for the Paired Loss Adaptation Baseline to 4.37\%. EER remains within 4.31--4.79\% under unseen prompts. Cross-domain evaluation on CN-Celeb yields a similar EER to the baseline, but lower accuracy at the native decision threshold. These findings support two-stage adaptation for improving speaker verification on the evaluated backbone.
Problem

Research questions and friction points this paper is trying to address.

Audio Large Language Models
Speaker Verification
Speaker Identity
Speaker-Preserving Representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-Stage Speaker Preservation
Audio Large Language Models
Low-Rank Adaptation
Speaker Verification
Parameter-Efficient
🔎 Similar Papers
No similar papers found.