Investigating the Role of Reasoning-Language Alignment in Monolingual Retrieval-Augmented Generation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how the alignment between the reasoning language and the languages of queries and documents affects model performance in retrieval-augmented generation (RAG). Methodologically, we construct an open-source German RAG benchmark based on *The Dark Eye* and employ an agent-based architecture to implement multilingual forced-reasoning strategies alongside cross-lingual ablation experiments. Results indicate that in long-context RAG scenarios, language alignment yields greater performance gains than raw language proficiency alone, with such benefits contingent upon context richness. Furthermore, German reasoning outperforms French yet remains inferior to English. These findings expose existing cross-lingual bottlenecks and underscore the necessity of developing native multilingual reasoning capabilities for RAG systems.
📝 Abstract
Reasoning traces improve large language models (LLMs), but current models are trained to reason mostly in English. It has been shown that forcing a model to reason in another language degrades accuracy, even when the reasoning language matches the language of the prompt -- but only for a setting where the model reasons over a short prompt. Here, we ask whether the same holds for retrieval-augmented generation (RAG), where the model must read and integrate a large amount of retrieved evidence in the target language. To study this, we build a fully monolingual German RAG question-answering testbed over the fictional world of the tabletop role-playing game The Dark Eye, a domain that is richly documented in German but too niche for the model to answer from memory, so that it has to rely on retrieval. Varying the forced reasoning language of an agentic RAG system on this testbed, we find that aligning the reasoning language with the language of the query and the retrieved documents helps. Forced German reasoning outperforms forced French, although the model benchmarks higher in French, so the benefit comes from alignment and not from language proficiency. The advantage grows when the retrieved context is richer and structure-aware. However, forced German only reaches the level of the model's native, unconstrained English reasoning without surpassing it, showing that native multilingual reasoning is needed. We publicly release the testbed and QA benchmark.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
Reasoning-Language Alignment
Multilingual Reasoning
Large Language Models
Monolingual RAG
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Generation
Reasoning-Language Alignment
Monolingual RAG
Agentic RAG
Multilingual Reasoning
🔎 Similar Papers
No similar papers found.