Score
Designs and implements systems that assign spoken or written utterances to individual speakers or characters by integrating audio, visual, and textual cues. Builds models and pipelines for multimodal speaker recognition and attribution tasks, including speaker diarization, cross‑modal alignment of transcripts with audio/video, and identity matching using voice and face/lip‑motion signals.
Traditional audio-based speaker diarization (SD) systems suffer from performance degradation under poor audio quality and high speaker voice similarity. This paper introduces the first purely text-driven SD paradigm, focusing exclusively on sentence-level speaker change detection and eliminating audio dependency entirely. Methodologically, we propose a Multi-Prediction Model (MPM) that synergistically integrates sequence labeling capabilities of pretrained language models, dialogue structure modeling, and multi-perspective semantic reasoning to enhance robustness and consistency in detecting speaker turns within short dialogues. Experiments across multiple dialogue datasets demonstrate that MPM significantly outperforms state-of-the-art audio-based methods, achieving a 12.3% absolute F1-score improvement in short-dialogue scenarios. To our knowledge, this is the first work to empirically validate the sufficiency and superiority of textual semantic features for speaker diarization, establishing text as a viable and effective modality for this task.
This work addresses the limited robustness and generalization of multimodal speaker recognition under modality-missing and cross-lingual conditions. To overcome the reliance of conventional approaches on complete modalities and monolingual assumptions, the study establishes the first systematic evaluation framework that integrates multimodal fusion, cross-lingual representation learning, and modality robustness modeling to effectively handle heterogeneous and incomplete inputs in non-ideal scenarios. The project delivers a reproducible benchmark platform and, for the first time in a unified challenge, incorporates settings with missing modalities and multilingual speakers. This setup reveals critical performance bottlenecks of current methods in realistic, complex environments and provides clear directions for future improvements.
This study investigates the effectiveness and robustness of automatic speech recognition (ASR) transcripts for speaker attribution, particularly examining the impact of transcription errors. Contrary to the conventional assumption that ASR errors degrade performance, experiments reveal that word-level errors do not significantly impair attribution accuracy—and may even introduce speaker-discriminative linguistic cues; in certain settings, ASR-generated transcripts yield superior attribution compared to human transcriptions. The work systematically evaluates how ASR system characteristics—including recognition accuracy, vocabulary coverage, and contextual modeling capability—affect attribution performance, corroborating findings through linguistic pattern analysis and speaker特征 modeling. Results demonstrate that speaker attribution based on ASR output is highly resilient and practically viable, challenging the prevailing assumption that high-fidelity transcription is a prerequisite. This establishes a novel paradigm for speaker identification in audio-deprived scenarios, where only ASR transcripts are available.
This work investigates the feasibility of authorship attribution models for speaker identification in speech transcription texts. Unlike written text, transcriptions lack punctuation and capitalization but contain speech-specific patterns such as fillers and backchannels. To address this, we introduce the first benchmark for speaker attribution in manually transcribed dialogues and propose a topic-controlled verification paradigm to mitigate topic confounding bias. Methodologically, we integrate contextual language models (BERT/RoBERTa) with n-gram and stylometric features, incorporating transcription-style analysis and domain-adaptive fine-tuning on speech-transcribed text. Experiments show that general-purpose models achieve moderate speaker discrimination under relaxed settings, but performance drops substantially under topic control. Crucially, fine-tuning on speech-transcribed corpora significantly improves speaker identification accuracy, demonstrating that speech-style features—e.g., disfluencies and interactional cues—are discriminative and recoverable from transcriptions.
To address the inconsistency and scene mismatch arising from the decoupled treatment of speaker extraction and diarization in complex overlapping speech, this paper proposes the first end-to-end jointly optimized framework that unifies frequency-domain speech separation with time-domain speaker activity annotation. The method integrates deep clustering, mask estimation, speaker activity detection, and waveform-level separation modules, supporting variable numbers of speakers and arbitrary overlap ratios. A bidirectional协同 mechanism enables mutual enhancement between extraction and diarization, breaking away from conventional cascaded pipelines. Evaluated on LibriMix, SparseLibriMix, and the real-world telephone conversation dataset CALLHOME, the approach achieves significant improvements over state-of-the-art methods on both tasks—marking the first demonstration of simultaneous gains in extraction quality (e.g., SI-SNRi) and diarization accuracy (e.g., DER).
Existing audio large language models struggle to achieve fine-grained understanding of speaker identity, vocal characteristics, and recording conditions, limiting their capacity for personalized and context-aware interaction. This work proposes SpeakerLLM, a unified framework that integrates a hierarchical speaker tokenizer—combining utterance-level embeddings with frame-level acoustic features—with a verification-oriented inference objective and a natural language interface. The model supports single-utterance profiling, pairwise speaker comparison, and evidence-based explainable reasoning by decoupling profiling evidence from final judgments, thereby generating structured decision trajectories. Experiments demonstrate that SpeakerLLM-Base outperforms general-purpose models in speaker and recording condition comprehension, while SpeakerLLM-VR maintains high accuracy and produces explanations aligned with supervised reasoning paradigms.
This study investigates whether incorporating visual information into multimodal speech recognition models exacerbates gender and racial biases. To this end, the authors introduce a controlled evaluation framework that pairs identical audio clips with synthetic videos of faces varying in gender and race, enabling systematic assessment of transcription performance across demographic groups. Using this setup, they evaluate prominent multimodal large language models, including mWhisper-Flamingo and Gemini, and uncover significant disparities—up to 4.05 percentage points in word error rate (WER)—between different demographic subgroups. These findings demonstrate that the visual modality can introduce new fairness risks in multimodal systems, offering critical empirical evidence to inform future bias evaluation and mitigation strategies in this domain.
This work addresses the challenge of accurately determining “who said what and when” in multi-person conversations, a task often compromised by visual biases and insufficient cross-modal alignment in existing methods. To overcome this, the authors propose the HumanOmni-Speaker model coupled with the VR-SDR task paradigm, which enables end-to-end spatiotemporal speaker grounding through natural language queries while rigorously avoiding visual shortcuts. A key innovation is the novel Visual Delta Encoder, which efficiently compresses inter-frame motion residuals from 25 fps video into just six tokens per frame, effectively capturing subtle lip movements and speaker trajectories. The approach further integrates high-frame-rate sampling, uncropped lip reading, and spatial localization. Experiments demonstrate state-of-the-art performance across multiple speaker-centric tasks, significantly advancing multimodal coordination and spatiotemporal localization accuracy.
This work addresses the challenge of speaker recognition in short audio clips from long-form TV dramas, where weak acoustic features often hinder performance. To tackle this issue, the authors propose DramaSR-LRM, a novel approach that introduces large reasoning models (LRMs) to the task for the first time. The method dynamically aggregates contextual evidence by autonomously invoking multimodal tools to fuse auditory, linguistic, and visual cues. Alongside the proposed method, the authors construct and release DramaSR-532K, a large-scale multimodal benchmark dataset comprising 532,000 character-annotated dialogue segments. Experimental results demonstrate that DramaSR-LRM substantially outperforms existing approaches, particularly excelling in scenarios with limited acoustic information.
Existing speaker diarization methods struggle with the challenges posed by open-domain scenarios such as films and TV shows, where speaker counts are large, audio-visual streams are often asynchronous, and environmental conditions are highly complex. This work proposes CineSRD, a novel framework that unifies visual, acoustic, and linguistic cues from video, speech, and subtitles for the first time. It leverages visual anchor clustering to register both on-screen and off-screen speakers and integrates an audio-language model to detect speaking turns. The authors construct and publicly release the first bilingual (Chinese–English) speaker diarization benchmark dataset for cinematic content. Extensive experiments demonstrate that CineSRD achieves state-of-the-art performance on this new dataset and remains competitive on conventional benchmarks, confirming its robustness and generalization capability in open-domain settings.