COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages
This study addresses the over-reliance on English pivoting and the scarcity of high-quality parallel corpora and evaluation benchmarks in machine translation for Indian languages. Departing from English-mediated approaches, we construct a parallel corpus comprising 1.16 million sentence pairs across 20 language pairs and multiple domains, sourced directly from native Indic texts. Furthermore, we introduce an expert-validated, domain-centric evaluation benchmark. Fine-tuning experiments on models such as IndicTrans2-Distilled and NLLB-200 demonstrate consistent performance improvements under both automatic and human evaluations. This work validates the effectiveness of high-quality native supervision and provides critical resources for advancing multilingual machine translation.