Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of protecting opt-out speaker privacy while accurately transcribing normal speech in multi-speaker automatic speech recognition (ASR). To this end, it proposes the first Target Speaker Unlearning ASR (TSU-ASR) framework. Specifically, a lightweight Enrollment-Conditioned Gating (ECG) module is designed and attached to a frozen dual-stream speech large language model. Integrated with end-to-end speaker diarization, this architecture enables the dynamic exclusion of unseen opt-out speakers during inference. Experimental evaluations on the AMI and AliMeeting datasets demonstrate that the proposed method significantly reduces transcription accuracy for opt-out speakers while maintaining stable error rates for non-opt-out participants. Ultimately, this work achieves dynamic privacy control without requiring model retraining.
📝 Abstract
We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires an ASR system to transcribe all speakers except the opt-out ones, while still indicating when those speakers are active. As a first step towards tackling this task, we introduce a novel, light-weight Enrollment-Conditioned Gating (ECG) module attachable to a frozen dual-stream speech LLM that enables ASR for new opt-out speakers dynamically during inference, even those who were not seen during initial ECG training phase. Our experiments on both AMI (English) and AliMeeting (Mandarin) datasets show that speech transcription accuracy for corresponding opt-out words or characters falls from 72.3% to 48.2% and from 73.6% to 27.3%, respectively, while retained speakers' transcription error rates maintain more or less the same. Our approach provides a practical solution for modern video conferencing platforms, allowing speakers to dynamically opt-out from automated AI transcriptions without forcefully leaving the meeting sessions, enabling a privacy-preserving interface for potentially millions of online meetings daily.
Problem

Research questions and friction points this paper is trying to address.

Target Speaker Unlearning
Automatic Speech Recognition
Privacy Preservation
Multi-speaker ASR
Speech Diarization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Target Speaker Unlearning
Enrollment-Conditioned Gating
Speech LLM
Inference-Time Adaptation
Privacy-Preserving ASR
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
B
Bo Su
Indiana University, Bloomington, USA
Y
Yueru Yan
Indiana University, Bloomington, USA
Thai Le
Thai Le
Assistant Professor in Computer Science, Indiana University
Machine learning & AI