🤖 AI Summary
In retrieval-augmented generation (RAG), decoder-only large language models exhibit strong positional sensitivity to retrieved passages—particularly overlooking critical information located in middle positions and suffering severe interference from irrelevant content (“middle-loss” problem). To address this, we propose a position-invariant RAG framework centered on a novel parallel key-value cache fusion mechanism. This mechanism comprises three components: (1) Transformer key-value cache reconstruction, (2) context-order-agnostic attention-based fusion, and (3) parallel cache injection—collectively decoupling model outputs from the positional ordering of retrieved content. Evaluated on three open-domain question answering benchmarks, our approach achieves strict positional invariance: error rates decrease by up to 37% compared to state-of-the-art RAG methods, while significantly improving robustness against irrelevant distractors and generation consistency.
📝 Abstract
Recent advancements in Large Language Models (LLMs) underscore the necessity of Retrieval Augmented Generation (RAG) to leverage external information. However, LLMs are sensitive to the position of relevant information within contexts and tend to generate incorrect responses when such information is placed in the middle, known as `Lost in the Middle' phenomenon. In this paper, we introduce a framework that generates consistent outputs for decoder-only models, irrespective of the input context order. Experimental results for three open domain question answering tasks demonstrate position invariance, where the model is not sensitive to input context order, and superior robustness to irrelevent passages compared to prevailing approaches for RAG pipelines.