StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reasoning challenges in long conversations arising from scattered evidence, shifting user preferences, and data scarcity. To this end, it proposes a cross-session retrieval and compositional reasoning framework based on tree-structured path tracing. Methodologically, the work introduces a sparse data-driven auxiliary task generation mechanism, combined with curriculum reinforcement learning and multi-session data augmentation, to effectively generalize short-context training to long-sequence scenarios. Experimental results demonstrate that a 14B-parameter model achieves 59% accuracy on the LongMemEval benchmark, surpassing a 32B baseline model and significantly enhancing long-range reasoning performance.
📝 Abstract
Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records across sessions and comparing timestamps to resolve branches, then recover the hidden target question among distractor leaves. We apply curriculum RL training progressively increasing tree depth and introduce a compositional variant whose edges carry step-level reasoning fragments, training the model to compose partial cues into coherent queries. Trained on 10K-token contexts, StateTree generalizes to 128K tokens without full-length RL costs and exhibits capabilities including cross-session retrieval, temporal reasoning, knowledge update, and compositional multi-hop reasoning. StateTree outperforms both SFT and RL-based baselines while preserving short-context general reasoning. StateTree-7B achieves gains up to +23.60% on LongMemEval (128k), and StateTree-14B reaches 59.00% accuracy on LongMemEval, surpassing QwenLong-L1-32B (45.20%).
Problem

Research questions and friction points this paper is trying to address.

long-term dialogue reasoning
cross-session retrieval
temporal reasoning
knowledge update
compositional multi-hop reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Long-Term Dialogue Reasoning
Curriculum Learning
Tree-Structured Path-Tracing
Length Generalization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Naen Xu
Zhejiang University
W
Wanqing Cui
Taobao & Tmall Group of Alibaba
Yibo Hu
Yibo Hu
NIO; previously JD AI Research/CASIA
Computer VisionFace AnalysisNASKnowledge DistillationMetric Learning
S
Shixin Hong
Taobao & Tmall Group of Alibaba
H
Hengyu An
Zhejiang University
M
Meiguang Jin
Taobao & Tmall Group of Alibaba
Junfeng Ma
Junfeng Ma
Mississippi State University
Design and ManufacturingLogisticsAI/MLHuman-Technology InteractionSustainability
Tianyu Du
Tianyu Du
Zhejiang University
AI SecurityAdversarial Machine Learning