Parallel Key-Value Cache Fusion for Position Invariant RAG

📅 2025-01-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In retrieval-augmented generation (RAG), decoder-only large language models exhibit strong positional sensitivity to retrieved passages—particularly overlooking critical information located in middle positions and suffering severe interference from irrelevant content (“middle-loss” problem). To address this, we propose a position-invariant RAG framework centered on a novel parallel key-value cache fusion mechanism. This mechanism comprises three components: (1) Transformer key-value cache reconstruction, (2) context-order-agnostic attention-based fusion, and (3) parallel cache injection—collectively decoupling model outputs from the positional ordering of retrieved content. Evaluated on three open-domain question answering benchmarks, our approach achieves strict positional invariance: error rates decrease by up to 37% compared to state-of-the-art RAG methods, while significantly improving robustness against irrelevant distractors and generation consistency.

Technology Category

Search and Optimization: Metareasoning and MetaheuristicsMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Image and Video Retrieval

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Recent advancements in Large Language Models (LLMs) underscore the necessity of Retrieval Augmented Generation (RAG) to leverage external information. However, LLMs are sensitive to the position of relevant information within contexts and tend to generate incorrect responses when such information is placed in the middle, known as `Lost in the Middle' phenomenon. In this paper, we introduce a framework that generates consistent outputs for decoder-only models, irrespective of the input context order. Experimental results for three open domain question answering tasks demonstrate position invariance, where the model is not sensitive to input context order, and superior robustness to irrelevent passages compared to prevailing approaches for RAG pipelines.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Retrieval-Augmented Generation
Position Sensitivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel Key-Value Cache Fusion
Retrieval Augmented Generation
Open-domain Question Answering
🔎 Similar Papers
2024-09-14arXiv.orgCitations: 6