π€ AI Summary
This study addresses the core challenge of βwho spoke what and whenβ in multilingual speaker-attributed automatic speech recognition (ASR) by systematically comparing cascaded and unified modeling strategies. Methodologically, a cascaded system comprising DiariZen, Qwen3-ASR, and an LLM-based generative error correction module is constructed, while VibeVoice is fine-tuned as a unified model baseline for comparison. To ensure reliable validation, only official data are utilized throughout the experiments. Results demonstrate that the cascaded approach exhibits superior robustness and performance on Task 1, whereas the unified speech large language model reveals significant potential for future development in this domain. Overall, this work provides empirical evidence to inform architectural design choices for multilingual speech understanding systems.
π Abstract
This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM). We address multilingual speaker-attributed ASR, where systems must determine who spoke when and what was spoken. We investigate two modeling strategies for this problem: a cascaded pipeline that combines speaker diarization with speech-LLM-based ASR, and a unified speech LLM that directly generates speaker labels, timestamps, and transcriptions. Our final submission is based on the cascaded pipeline, consisting of a fine-tuned DiariZen diarization model, a fine-tuned Qwen3-ASR model, and LLM-based generative error correction. For comparison, we also fine-tune VibeVoice-ASR as a unified model using the same official training data. All task-specific fine-tuning and model selection are performed using only the official MLC-SLM data, without external data or pseudo-labels. Experimental results demonstrate that the cascaded system remains more reliable under the MLC-SLM Task 1 conditions, while unified speech LLMs offer a promising direction for future speaker-attributed ASR.