Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of unbounded historical memory growth and the absence of context-aware thought activation mechanisms in intelligent agents. To this end, we propose a Context-Conditioned Thought Memory framework that transcends conventional retrieval-compression paradigms. Specifically, it employs a lightweight thought policy network to internalize cross-cycle reasoning experiences into predictive strategies, thereby enabling spatiotemporal evolution-based long-range pattern discovery and dynamic knowledge integration. Experimental results demonstrate that the proposed framework achieves perfect scores of 1.0 in both temporal rule generalization F1 and relation discovery accuracy. Furthermore, it yields substantial improvements in overall reasoning performance while reducing online processing latency by approximately 90%.
📝 Abstract
Long-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely. However, existing memory mechanisms mainly retrieve, summarize, or compress past content and do not directly learn when particular kinds of thinking should be activated or discover new thinking knowledge from temporally dispersed experiences. This paper proposes a situation-conditioned thinking memory framework that transforms historical reasoning experience into a lightweight policy for predicting what should be thought about in the current situation, while leaving detailed reasoning to a large language model. Situations may represent temporal or spatiotemporal evolution rather than only current states. Temporary experiences are also periodically analyzed across multiple independent episodes to identify repeated long-range regularities, which are consolidated into new thinking knowledge and further internalized by the lightweight policy. Experiments show that the learned policy achieves 1.000 F1 on temporal-rule generalization, improves DeepSeek reasoning F1 from 0.789 to 0.868, reduces online processing time from 0.3636 ms to 0.0382 ms per query at 30,000 historical situations, and reaches 1.000 relation-discovery F1 and future-thinking accuracy after sufficient repeated cross-experience evidence.
Problem

Research questions and friction points this paper is trying to address.

Long-term LLM Agents
Reasoning Experience Reuse
Memory Mechanisms
Situation-Conditioned Thinking
Knowledge Discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Situation-Conditioned Thinking Policy
Long-Term LLM Agents
Lightweight Memory Framework
Cross-Episode Knowledge Consolidation
Temporal Reasoning
🔎 Similar Papers
No similar papers found.