Learning to Accumulate Knowledge with Mutual Information

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of redundant knowledge accumulation and insufficient contribution from novel experiences in LLM-based agents by proposing the Knowledge Weaver framework. Leveraging Group Relative Policy Optimization (GRPO), this method trains a language model as a curator through token-level mutual information feedback and marginal success rewards, enabling it to extract reusable knowledge that is both distinctive and effective from interaction trajectories. Experimental results demonstrate that the proposed framework achieves success rates of 54.0% on ALFWorld and 42.0% on WebShop, significantly outperforming existing baselines and human-curated knowledge bases. Overall, this work establishes an efficient reinforcement learning paradigm for agent knowledge management.
📝 Abstract
Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions. However, curating new experiences into a knowledge bank that becomes more useful as it grows remains challenging. Effective knowledge accumulation should limit redundant overlap among entries and ensure that new knowledge contributes beyond what the bank already provides. Yet training a curator with Group Relative Policy Optimization (GRPO) on standalone task success can reinforce general guidance even when it duplicates existing knowledge. Therefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories. We couple feedback inspired by token-wise mutual information (MI) with marginal success rewards to guide knowledge accumulation. Together, these signals encourage the curator to preserve distinct information from experience and produce entries that improve task success when added to existing knowledge. Standalone success rewards also favor entries that are useful on their own. On ALFWorld and WebShop, Knowledge Weaver achieves mean success rates of 54.0\% and 42.0\% with k=10 retrieved entries, exceeding GRPO by 16.9 and 18.7 percentage points, respectively. Its knowledge banks also outperform the evaluated prompt-based and established banks, including human-written banks, in overall ALFWorld success rate and WebShop score with the executor frozen. Our codebase is available at https://github.com/LaoKuiZe/Knowledge-Weaver.
Problem

Research questions and friction points this paper is trying to address.

knowledge accumulation
large language model agents
redundancy reduction
knowledge curation
mutual information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mutual Information
Reinforcement Learning
Knowledge Accumulation
Large Language Model Agents
Knowledge Curation
🔎 Similar Papers