FASTDIAR: Frame-level speaker encoder for Streaming Diarization

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high computational overhead and performance degradation in multi-speaker scenarios associated with streaming speaker diarization for real-time conversational agents. To this end, it proposes a low-latency, CPU-friendly streaming diarization system. Methodologically, the work reformulates the speaker recognition architecture into a causal frame-level encoder and introduces a novel knowledge distillation paradigm from the utterance level to the frame level, coupled with an online clustering algorithm utilizing self-similarity gated updates. Experimental results demonstrate that the proposed system achieves state-of-the-art accuracy with sub-second latency while performing inference at five times real-time speed on a single CPU core. It significantly outperforms cache-based baselines and effectively mitigates performance degradation in multi-speaker settings.
πŸ“ Abstract
Real-time conversational agents require speaker diarization that streams and runs on a CPU. Most systems apply an utterance-level speaker encoder to short, heavily overlapping chunks, which wastes computation and leaves the model optimized for the wrong task. We instead turn a state-of-the-art speaker recognition architecture into a causal frame-level encoder that reads the stream once and emits one embedding every 80~ms from a bounded two-second window of past audio, and pair it with online clustering that gates every update on the self-similarity of the stream. Trained only by distillation from an utterance-level teacher on simulated and out-of-domain mixtures, and evaluated with one fixed set of hyperparameters, the system is the most accurate streaming diarizer on low-overlap benchmarks at sub-second latency, degrades far less than cache-based systems as the number of speakers grows, and runs five times faster than real time on a single CPU thread.
Problem

Research questions and friction points this paper is trying to address.

speaker diarization
streaming
real-time
frame-level speaker encoder
CPU
Innovation

Methods, ideas, or system contributions that make the work stand out.

frame-level speaker encoder
streaming diarization
causal architecture
knowledge distillation
online clustering
πŸ”Ž Similar Papers
No similar papers found.