COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the over-reliance on English pivoting and the scarcity of high-quality parallel corpora and evaluation benchmarks in machine translation for Indian languages. Departing from English-mediated approaches, we construct a parallel corpus comprising 1.16 million sentence pairs across 20 language pairs and multiple domains, sourced directly from native Indic texts. Furthermore, we introduce an expert-validated, domain-centric evaluation benchmark. Fine-tuning experiments on models such as IndicTrans2-Distilled and NLLB-200 demonstrate consistent performance improvements under both automatic and human evaluations. This work validates the effectiveness of high-quality native supervision and provides critical resources for advancing multilingual machine translation.
📝 Abstract
Machine translation (MT) for Indian languages remains constrained by the limited availability of high-quality, Indic-centric parallel corpora and evaluation benchmarks. Existing multilingual resources are largely constructed from English-pivot content and often fail to capture the linguistic diversity, cultural complexity, and domain-specific characteristics of Indian languages. We present COILD, an Indic-centric parallel corpus comprising over 1.16 million human-translated and human-verified sentence pairs, covering 20 Indian language pairs across the Indo-Aryan, Dravidian, Tibeto-Burman, and Austro-Asiatic language families. The corpus is built entirely from original Indian language sources collected from licensed repositories spanning eight domains with direct real-world applicability. Furthermore, we introduce a domain-centric benchmark comprising 2,000 expert-verified sentences to enable consistent multilingual and cross-lingual evaluation across Indian language pairs. To validate the effectiveness of COILD, we fine-tune two representative multilingual neural machine translation models, IndicTrans2-Distilled and NLLB-200. Experimental results demonstrate consistent improvements across language pairs, domains, automatic evaluation metrics, and human evaluation, highlighting the effectiveness of high-quality Indic-centric supervision. COILD provides a valuable training and evaluation resource for advancing multilingual machine translation and future multilingual language models for Indian languages.
Problem

Research questions and friction points this paper is trying to address.

Machine Translation
Indian Languages
Parallel Corpus
Evaluation Benchmark
Low-resource
Innovation

Methods, ideas, or system contributions that make the work stand out.

Indic-centric parallel corpus
machine translation benchmark
multilingual NMT
low-resource languages
cross-lingual evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kshetrimayum Boynao Singh
IIT Patna
N
Nitin Kumar Mishra
IIT Delhi
P
Palash Pratim Dutta
IIT Guwahati
A
Atai Waris Khan
IIIT Delhi
A
Aparna Kaushik
IGDTUW
Avinash Kumar
Avinash Kumar
Research Assistant Soongsil University, Seoul, South Korea
Machine LearningDeep LearningComputer Vision GAN's
D
Deeksha
MIT-MAHE
D
Deepak Kumar
IIT Patna
S
Saroj Kumar Jha
IIT Patna
S
Saloka Sengupta
IIT Patna
A
Anansa Roy
IIT Patna
U
Umalatha Kannoth
MIT-MAHE
S
Saifulla Samar
IIT Guwahati
M
Meena Sharma
IIIT Delhi
M
Manpreet Kaur
IGDTUW
Jyoti Sharma
Jyoti Sharma
IIT Delhi
A
Ashwini Vaidya
IIT Delhi
M
Muralikrishna SN
MIT-MAHE
Md Shad Akhtar
Md Shad Akhtar
IIIT Delhi
NLPConversational DialogMental-HealthMisinformationCode-mixed Languages
Poonam Bansal
Poonam Bansal
IGDTUW
A
Amita Dev
IGDTUW
Sanasam Ranbir Singh
Sanasam Ranbir Singh
Professor of Computer Science and Engineering, IIT Guwahati
Information retrievalMachine LearningData Mining
Samit Bhattacharya
Samit Bhattacharya
IIT Guwahati
Tanmoy Chakraborty
Tanmoy Chakraborty
Associate Professor, IIT Delhi, India
Natural Language ProcessingLarge Language ModelsSocial Computing
Asif Ekbal
Asif Ekbal
Department of Computer Science and Engineering, IIT Patna
Artificial IntelligenceNatural Language ProcessingMachine Learning Application