Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing instruction adherence with translation quality in instruction-following machine translation. To this end, it proposes a two-stage data curation pipeline and introduces the ChindaMT model family. Methodologically, this work pioneers a reference-based constraint extraction technique to ensure instruction feasibility, and constructs a high-quality dataset by integrating IFD score filtering, constraint-based filtering, and large language model fine-tuning. Experimental results demonstrate that ChindaMT consistently outperforms baseline models across multiple parameter scales, achieving a maximum win rate of 68.4%. The model weights and datasets have been made publicly available.
📝 Abstract
Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.
Problem

Research questions and friction points this paper is trying to address.

Instruction-Following Machine Translation
Rule Compliance
Translation Quality
Data Curation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reference-Grounded Data Curation
Instruction-Following Machine Translation
Instruction-Following Difficulty (IFD)
Constraint Extraction
Data Augmentation
🔎 Similar Papers
No similar papers found.