🤖 AI Summary
This study addresses the challenge of balancing instruction adherence with translation quality in instruction-following machine translation. To this end, it proposes a two-stage data curation pipeline and introduces the ChindaMT model family. Methodologically, this work pioneers a reference-based constraint extraction technique to ensure instruction feasibility, and constructs a high-quality dataset by integrating IFD score filtering, constraint-based filtering, and large language model fine-tuning. Experimental results demonstrate that ChindaMT consistently outperforms baseline models across multiple parameter scales, achieving a maximum win rate of 68.4%. The model weights and datasets have been made publicly available.
📝 Abstract
Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.