Agri-Query: A Case Study on RAG vs. Long-Context LLMs for Cross-Lingual Technical Question Answering

📅 2025-08-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study evaluates large language models (LLMs) on cross-lingual technical question answering, specifically using multilingual agricultural machinery user manuals (English/French/German): English queries are answered via cross-lingual retrieval and generation. To stress-test robustness, we introduce “needle-in-a-haystack” long-context challenges (128K tokens) and unanswerable questions to assess hallucination. We propose a hybrid retrieval-augmented generation (RAG) framework integrating keyword-based retrieval, semantic search, and LLM-as-a-judge evaluation—significantly improving robustness over pure long-context baselines. Experiments show >85% accuracy across all three languages, with Gemini 2.5 Flash and Qwen 2.5 7B achieving top performance. Our method effectively mitigates hallucination, enhances cross-lingual retrieval precision, and provides an open-source, reproducible industrial-grade evaluation framework. The work empirically reveals critical trade-offs between RAG and long-context approaches in real-world technical documentation QA.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Question AnsweringData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
We present a case study evaluating large language models (LLMs) with 128K-token context windows on a technical question answering (QA) task. Our benchmark is built on a user manual for an agricultural machine, available in English, French, and German. It simulates a cross-lingual information retrieval scenario where questions are posed in English against all three language versions of the manual. The evaluation focuses on realistic "needle-in-a-haystack" challenges and includes unanswerable questions to test for hallucinations. We compare nine long-context LLMs using direct prompting against three Retrieval-Augmented Generation (RAG) strategies (keyword, semantic, hybrid), with an LLM-as-a-judge for evaluation. Our findings for this specific manual show that Hybrid RAG consistently outperforms direct long-context prompting. Models like Gemini 2.5 Flash and the smaller Qwen 2.5 7B achieve high accuracy (over 85%) across all languages with RAG. This paper contributes a detailed analysis of LLM performance in a specialized industrial domain and an open framework for similar evaluations, highlighting practical trade-offs and challenges.
Problem

Research questions and friction points this paper is trying to address.

Evaluating long-context LLMs vs RAG for technical QA
Testing cross-lingual information retrieval in agricultural manuals
Assessing model performance on needle-in-haystack challenges
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid RAG outperforms long-context prompting
Gemini and Qwen models achieve high accuracy
Framework for cross-lingual technical QA evaluation
J
Julius Gun
Chair of Agrimechatronics, Technical University of Munich (TUM), Munich, Germany
Timo Oksanen
Timo Oksanen
Technical University of Munchen (TUM)
automationcontrol engineeringroboticsmechatronicstractors