Contract Memory Compiler: Resolve, Then Traverse

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of erroneous evidence selection in multi-hop question answering caused by information updates within long-horizon external memory. To mitigate this, we propose CMC, a method employing a "parse-then-traverse" strategy. Specifically, it first leverages large language models to extract and update relational graphs for parsing the current state, and subsequently traces relevant raw records dynamically. This provides the answering model with accurate evidence in a single pass, ensuring that retrieval relies on dynamic states rather than static matching. Experimental results demonstrate that CMC achieves state-of-the-art performance with 78.25% multi-hop accuracy on the FactConsolidation dataset. Furthermore, we introduce MQuAKE-MemStream, a new benchmark designed to facilitate future research in this domain.
📝 Abstract
External memory lets language-model agents answer questions about histories too long for the answer model's context window. Updates create a harder problem than retrieving a recent fact: changing one relation can redirect a multi-hop question to records about an entity absent from the question. We study this update-dependent evidence selection problem and introduce the Contract Memory Compiler (CMC). Before seeing a question, CMC uses a language model to identify relations in the history and record where each one was stated. It applies later updates to determine the current relations, follows them from entities named in the question, and passes the corresponding original records to the answer model in one call. Thus the current state determines which evidence is read, rather than merely refreshing values in a previously selected context. To the best of our knowledge, CMC achieves state-of-the-art multi-hop accuracy on FactConsolidation, reaching 78.25% overall and 61.0% at 262K. With the extracted relations and answer model held fixed, selecting evidence before resolving updates reduces multi-hop accuracy to 21.50%. We also introduce MQuAKE-MemStream, a derived dataset of ordered memory streams built from MQuAKE-Remastered counterfactual cases.
Problem

Research questions and friction points this paper is trying to address.

external memory
multi-hop question answering
evidence selection
memory updates
language-model agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contract Memory Compiler
Multi-hop reasoning
Evidence selection
Memory updates
Language model agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhi Song
Department of Computer Science, City University of Hong Kong, Hong Kong SAR
X
XiMing Xing
Tencent, China
C
Chunhan Li
Tencent, China
Weian Mao
Weian Mao
MIT CSAIL
Z
Zhenchao Tang
Tencent, China
H
Hanbo Huang
Tencent, China
F
Fan Xu
Tencent, China
Jiale Zhou
Jiale Zhou
MDH
requirements engineeringsafety critical systemshazard analysisontology
J
Jiahui Guan
Tencent, China
Z
Zejian Ding
Department of Computer Science, City University of Hong Kong, Hong Kong SAR
Chen Ma
Chen Ma
Assistant Professor, City University of Hong Kong
Recommender SystemsData MiningData-Centric AISocial Computing
L
Lusheng Wang
Department of Computer Science, City University of Hong Kong, Hong Kong SAR