Fair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoSh

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant cross-lingual performance gap in multilingual fact-checking with large language models, where non-English performance substantially lags behind English. Our analysis reveals that factual knowledge is already encoded within the model but suffers from a "retrieval failure." To address this, we propose RoSh, a training-free method that employs linear probes to localize multilingual biases within the residual stream and computes language-specific offset and rotation matrices via closed-form solutions under orthogonality constraints, thereby rectifying the deficiency through geometric transformations. Experiments demonstrate that RoSh closes 75% of the performance gap on average, elevating Llama-3B's Arabic accuracy from chance level to near-English parity and outperforming existing baselines by a factor of five to thirteen.
📝 Abstract
Misinformation on social media remains a critical problem, and more and more people settle it by asking a language model instead of a fact checker. Whether models judge such claims reliably is debated; whether they judge them equally well in every language people ask in has gone almost unasked. We test eight models from five families, 3B to 70B, on 1,500 encyclopedic factual claims that exist in identical form in eight languages. English is judged better than every other language on every model, and the gap is widest on the smallest ones, where Llama-3B on Arabic is no better than guessing. Existing remedies retrain on more multilingual data or fit an unconstrained map between language representations, and neither asks whether the model already holds the answer and simply fails to say it. It largely does: a linear probe recovers the truth from the very activations the model fails to express. We propose RoSh, a per-language shift and rotation of the residual stream, computed in closed form at three layers, with no training and no weight modified. It improves every model and closes 75% of the gap on average, helping most where the model was worst: Arabic on Llama-3B goes from chance to nearly the English level, and a fifth fewer of the claims answered correctly in English are lost in translation. What remains is no longer a read-out failure: afterwards the head recovers as much of what is encoded outside English as it does in English. An unconstrained map fitted on the same pairs falls below the untouched baseline, so the orthogonality constraint is doing the work, and every model clears a scrambled-correspondence control and ten further controls. On the two benchmarks of the closest inference-time method, latent-space intervention, run with its own data and metric code, RoSh's gains are five to thirteen times larger.
Problem

Research questions and friction points this paper is trying to address.

Fact-Checking
Cross-Lingual Gap
Large Language Models
Misinformation
Fairness
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-lingual fact-checking
residual stream intervention
training-free alignment
orthogonal rotation
large language models
🔎 Similar Papers
2024-06-20International Conference on Computational LinguisticsCitations: 2