Improving Low-Resource Translation with Dictionary-Guided Fine-Tuning and RL: A Spanish-to-Wayuunaiki Study

📅 2025-08-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Machine translation for low-resource language pairs—specifically Spanish→Wayuunaiki—faces severe challenges due to insufficient pretraining data and scarce parallel corpora. Method: We propose a dictionary-guided reinforcement learning framework that treats an external bilingual dictionary as a callable tool, enabling the model to autonomously decide when to query it and how to integrate retrieved lexical information during decoding. The framework combines supervised instruction fine-tuning with guided reward policy optimization (GRPO), using BLEU score as a sparse reward signal for end-to-end training—without explicit dictionary injection or hand-crafted rules. Contribution/Results: Our approach significantly improves lexical and morphological generalization for low-resource languages. On the AmericasNLP 2025 test set, it achieves an absolute BLEU gain of 3.37 points (18% relative improvement) over a dictionary-free baseline, demonstrating the effectiveness and scalability of tool-augmented reinforcement learning in low-resource MT.

Technology Category

Natural Language Processing: Lexical Semantics and MorphologyMachine Learning: Reinforcement LearningPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Low-resource machine translation remains a significant challenge for large language models (LLMs), which often lack exposure to these languages during pretraining and have limited parallel data for fine-tuning. We propose a novel approach that enhances translation for low-resource languages by integrating an external dictionary tool and training models end-to-end using reinforcement learning, in addition to supervised fine-tuning. Focusing on the Spanish-Wayuunaiki language pair, we frame translation as a tool-augmented decision-making problem in which the model can selectively consult a bilingual dictionary during generation. Our method combines supervised instruction tuning with Guided Reward Policy Optimization (GRPO), enabling the model to learn both when and how to use the tool effectively. BLEU similarity scores are used as rewards to guide this learning process. Preliminary results show that our tool-augmented models achieve up to +3.37 BLEU improvement over previous work, and a 18% relative gain compared to a supervised baseline without dictionary access, on the Spanish-Wayuunaiki test set from the AmericasNLP 2025 Shared Task. We also conduct ablation studies to assess the effects of model architecture and training strategy, comparing Qwen2.5-0.5B-Instruct with other models such as LLaMA and a prior NLLB-based system. These findings highlight the promise of combining LLMs with external tools and the role of reinforcement learning in improving translation quality in low-resource language settings.
Problem

Research questions and friction points this paper is trying to address.

Enhancing low-resource machine translation with external dictionary integration
Addressing limited parallel data for Spanish-Wayuunaiki language pair
Improving translation quality through reinforcement learning and tool consultation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dictionary-guided fine-tuning for low-resource translation
Reinforcement learning with BLEU reward optimization
Tool-augmented decision-making during translation generation
🔎 Similar Papers
No similar papers found.
M
Manuel Mosquera
Universidad de los Andes, Bogotá, Colombia
M
Melissa Robles
Universidad de los Andes, Bogotá, Colombia
J
Johan Rodriguez
Universidad de los Andes, Bogotá, Colombia
Ruben Manrique
Ruben Manrique
Assistant Professor at Universidad de los Andes, Systems and Computing Engineering Department
Semantic WebArtificial IntelligenceNatural Language Processing