Voices as Handles: Reasoning about Speaker Identity with Frozen Text LLMs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of large language models (LLMs) in acoustic identity awareness, which hinders multi-user cross-session speaker attribution reasoning. To overcome this, we propose a Speaker Handles mechanism that maps voiceprints into soft tokens via a lightweight projector and injects them into a frozen LLM. Methodologically, a three-stage curriculum learning strategy is employed to train the projector for generating high-quality acoustic representations, while a novel benchmark, SpeakerBind, is constructed to bridge existing evaluation gaps. Experimental results demonstrate that this framework efficiently bridges the acoustic and semantic spaces with minimal parameters, achieving accuracies of 97.4% on VoxCeleb1 and 70.4% on the SpeakerBind benchmark, closely approaching theoretical upper bounds.
📝 Abstract
Multi-user voice agents must track who said what across dialogue sessions. Text LLMs are attractive backbones for such agents, but transcripts alone do not expose acoustic speaker identity, leaving the model without a persistent reference for linking information to speakers across sessions. We address this gap by introducing Speaker Handles, soft-token representations that expose acoustic speaker identity to a frozen text LLM for cross-session speaker-dependent reasoning. A three-stage curriculum trains a lightweight projector, with fewer than 0.1% of the backbone's parameters, to map speaker embeddings into these handles. Establishing whether the resulting handles truly support cross-session speaker-dependent reasoning is challenging with existing benchmarks because textual cues can partially reveal fact ownership. We therefore present SpeakerBind, a controlled shared-agent benchmark in which overlapping facts across users require correct cross-session speaker attribution. Speaker Handles achieve 97.40-98.36% accuracy on VoxCeleb1 and 70.40% on SpeakerBind, close to the 71.88% topline. These results show that the proposed Speaker Handles provide an efficient way to integrate acoustic speaker identity into frozen text LLMs for speaker-content reasoning.
Problem

Research questions and friction points this paper is trying to address.

speaker identity
cross-session reasoning
text LLMs
multi-user voice agents
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speaker Handles
Frozen Text LLMs
Cross-session Reasoning
Lightweight Projector
SpeakerBind Benchmark
🔎 Similar Papers