🤖 AI Summary
This study addresses the data association ambiguity in unordered direction-of-arrival (DOA) estimation caused by speech intermittency, spatial proximity, and complex acoustic environments. To this end, we propose an identity-assisted multi-speaker tracking method that innovatively fuses long-term stable speaker embeddings with short-term continuous spatial cues. A unified neural tracker is designed to map multi-source observations into consistent identity trajectories by leveraging a temporal self-attention module to capture trajectory evolution and a source attention mechanism to disambiguate competing tracks. Experimental results demonstrate that the proposed approach effectively mitigates association confusion under multi-source competition, significantly enhancing the reliability of speech source tracking in complex scenarios.
📝 Abstract
Tracking speech sources remains a challenge due to ambiguous data association arising from intermittent speech, close spatial proximity, and complex acoustic conditions. To address these issues, we propose an identity-assisted association that maps unordered direction-of-arrival (DOA) estimates to speaker-consistent source trajectories for reliable speech source tracking. Specifically, speaker identity embeddings are directly integrated into the model input as a complementary cue to spatial features. This enables maintaining identity consistency by combining long-term time-invariant vocal identity characteristics with the short-term continuity of spatial cues. To effectively process these heterogeneous inputs while accommodating their distinct characteristics, we design a unified neural tracker. Within this model, time self-attention modules capture the temporal evolution of each source, while source self-attention modules distinguish between competing source tracks. Experimental results demonstrate the superiority of the proposed neural tracker in mitigating association confusion for speech source tracking.