Role-aware Heuristic Episodic Attention for Conversational LLMs

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issues of instruction forgetting and information loss caused by context accumulation in multi-turn dialogues with large language models. To this end, it proposes the REA framework, which innovatively introduces a role-aware strategy to decouple persistent global instructions from dynamic episodic interaction memory. By integrating heuristic retrieval with text compression algorithms to optimize the representation of historical turns, REA achieves efficient hierarchical context management. Experimental results demonstrate that REA improves evaluation scores by 16.5% on the Long-MT-Bench+ benchmark while reducing inference latency by 2.91 times. Furthermore, its generalizability and effectiveness are validated across multiple parameter scales and both Chinese and English role-playing tasks.
📝 Abstract
Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attention), a context-management framework that assigns different persistence and representation policies to instructions and episodic interactions. Instructional Memory retains identified global constraints in a dedicated prefix. Episodic Memory preserves user inputs and compresses model replies, while heuristic retrieval selects raw text, compressed representations, or omission for each historical turn. On Long-MT-Bench+, REA improves the judge score from 6.32 to 7.36 on a 10-point scale, a 16.5% relative gain over the Vanilla baseline, and reduces average latency by 2.91$\times$. Additional evaluations show aggregate gains on three backbones spanning 1.7B-7B parameters and on Chinese and English role-playing tasks. These results support role-aware context management as a practical approach to maintaining conversational continuity and instruction adherence.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Multi-turn Conversations
Contextual Decay
Attention Pollution
Instruction Adherence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Role-aware Attention
Episodic Memory
Context Management
Heuristic Retrieval
Conversational LLMs
🔎 Similar Papers
No similar papers found.
W
Wanyang Hong
National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha, China
Zhaoning Zhang
Zhaoning Zhang
National University of Defense Technology
MLSysCompute VisionDistributed Computing
Y
Yi Chen
National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha, China
L
Libo Zhang
National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha, China
B
Baihui Liu
National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha, China
Linbo Qiao
Linbo Qiao
NUDT
Stochastic OptimizationDistributed OptimizationLarge-scale Machine Learning
Z
Zhiliang Tian
National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha, China
Dongsheng Li
Dongsheng Li
Professor, School of Computer Science, National University of Defense Technology
Distributed ComputingParallel ComputingCloud ComputingPeer-to-Peer ComputingBig Data