STAM-ASR: Speaker-Temporal Anchoring with Memory for Multi-Speaker ASR

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of multi-speaker speech recognition and attribution in natural conversations, where speaker alternation, overlap, and recurrence complicate processing. To this end, this work proposes a lightweight framework that eliminates the need for external diarization systems by extending pretrained AudioLLMs. Specifically, the approach leverages intermediate representations to learn speaker activity, introduces a temporal anchoring mechanism to provide explicit spatiotemporal cues, and designs a fixed-size memory network for cross-turn context modeling. Experiments on datasets such as AMI validate the complementary gains of spatiotemporal conditioning and the memory module. Furthermore, the findings highlight that robust speaker tracking remains a critical challenge requiring further investigation.
📝 Abstract
Natural conversations make both speech recognition and speaker attribution challenging for ASR, as speakers take turns, overlap, and reappear over time. We propose STAM-ASR, Speaker-Temporal Anchoring with Memory, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR. Without relying on an external diarization system, STAM-ASR learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features. Hence providing explicit who and when cues to modulate the AudioLLM's semantic representation without explicit speech separation. STAM-ASR further maintains fixed-size speaker and conversational memories to carry complementary context across turns. We evaluate STAM-ASR on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions. Our reported results shows that speaker-temporal conditioning and memory provide complementary benefits, while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
Problem

Research questions and friction points this paper is trying to address.

Multi-speaker ASR
Speaker attribution
Speech recognition
Speaker tracking
Overlapping speech
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Speaker ASR
Speaker-Temporal Anchoring
AudioLLM
Memory Mechanism
Speaker Diarization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
V
Victor Tolulope Olufemi
Qatar Computing Research Institute (QCRI), Doha, Qatar
S
Syeda Faiza Ahmed Sara
Qatar Computing Research Institute (QCRI), Doha, Qatar
Shammur Absar Chowdhury
Shammur Absar Chowdhury
Qatar Computing Research Institute
Conversational AIRepresentation LearningDeep LearningSpeech processingNLP