🤖 AI Summary
This study addresses the challenges of capturing the informality inherent in spoken dialogue and the absence of evaluation benchmarks for Hindi dialogue translation. To bridge this data gap, we construct the first large-scale multilingual corpus comprising 49,000 dialogues. Methodologically, we fine-tune open-source large language models on this corpus and conduct a multidimensional evaluation integrating automatic metrics, LLM-as-a-judge paradigms, and human Scalar Quality Metric (SQM) assessments. Experimental results demonstrate that the fine-tuned models significantly outperform baseline systems, effectively enhancing dialogue translation quality. By establishing a comprehensive benchmark and a robust evaluation framework, this work sets a new standard for spoken dialogue translation tasks and provides valuable resources for future research in low-resource and informal conversational machine translation.
📝 Abstract
Existing translation models are typically trained on sentence-level and formal text, limiting their ability to capture everyday conversational dialogue phenomena such as informality, speaker interaction, and discourse coherence. Most existing Indic translation resources and evaluation benchmarks focus on sentence-level or formal text, making it difficult to assess translation quality of the dialogue phenomena. In this work, we introduce BaatCheet, a multilingual dialogue corpus named after the Hindi term for conversation or chitchat, containing approximately 49,000 dialogues for dialogue translation across five translation directions. We fine-tune five open-source LLMs across seven training data configurations and find that fine-tuning yields substantial gains over zero- and few-shot baselines. To comprehensively evaluate dialogue translation quality, we employ multiple evaluation strategies, including automatic metrics, LLM-as-judge, and human assessments using an SQM-guided Direct Assessment (DA) Protocol.