Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the information bias in large language models (LLMs) caused by query-language preference during multilingual knowledge conflicts. To investigate this, we construct Waldo, a cross-lingual question-answering benchmark, and systematically evaluate eight mainstream LLMs. Attention mechanism analysis reveals that while models exhibit no significant language preference under knowledge absence, they demonstrate strong reliance on the query language during knowledge conflicts. Accordingly, we propose two intervention strategies—attention head ablation and LoRA-based parameter-efficient fine-tuning—to mitigate this bias. Experimental results show that our methods reduce the query-language preference gap by 61.5%, providing an effective solution for enhancing the knowledge consistency and fairness of multilingual LLMs.
📝 Abstract
Large Language Models increasingly serve as interfaces for knowledge-intensive information seeking tasks across languages by synthesizing multilingual evidence. Prior work has shown that they often exhibit query-language preference -- the tendency to favor sources written in the language of the query -- but has largely examined this behavior in settings where equivalent knowledge is available across languages. However, this bias becomes consequential when sources in different languages provide incomplete or inconsistent accounts of the same fact, since the information users receive then depends on the sources a model selects to use. To characterize query-language preference under such cross-lingual knowledge disparities, we introduce Waldo, a multilingual Question-Answering (QA) benchmark constructed from Wikipedia. Waldo contains 12K QA pairs targeting knowledge gaps, where a fact is available in one language but absent in another, and knowledge conflicts, where language editions provide conflicting versions of the same fact. Evaluating eight models across five languages, we find that when one language edition merely lacks the relevant fact, models generally use evidence from the other language regardless of the query language. Under conflicting accounts, however, model responses strongly align with the document in the query language, causing semantically equivalent queries to elicit different accounts depending on the user's language. Finally, we explore two different approaches that could mitigate this preference under knowledge conflicts: a mechanistic intervention that ablates attention heads associated with query-language preference, and LoRA-based training, which reduces the preference gap by up to 61.5%.
Problem

Research questions and friction points this paper is trying to address.

query-language preference
cross-lingual knowledge disparities
large language models
multilingual question answering
knowledge conflicts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Query-language Preference
Cross-lingual Knowledge Disparities
Multilingual QA Benchmark
Mechanistic Intervention
LoRA-based Training