🤖 AI Summary
This work addresses the critical scarcity of high-quality multilingual code-mixed dialogue data for Indic languages, particularly the lack of large-scale resources integrating English with native languages in both native scripts and Romanized forms. The paper introduces IndicTalk, a novel corpus comprising 1.32 million event-driven, role-conditioned multi-turn code-mixed dialogues spanning nine Indic language families and eighteen language variants—the first of its kind. Leveraging multilingual large language models, the approach generates dialogues conditioned on real-world news events as contextual anchors and incorporates an automated quality validation pipeline. Comprehensive evaluations—linguistic, automatic, and human—demonstrate that the corpus exhibits high naturalness, fluency, and linguistic diversity. The dataset is publicly released to significantly advance research in multilingual conversational AI for Indic languages.
📝 Abstract
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .