EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited transferability of LLM agents in unseen environments and the high training costs and error accumulation associated with existing world models. To overcome these challenges, this work proposes EVOKE, a post-training method that introduces goal diversity under fixed states and leverages multi-goal ranked contrastive learning to suppress superficial habit reliance, thereby eliciting latent world knowledge from pretraining to support transferable decision-making. Theoretically, it is proven that cross-goal capabilities implicitly encode a recoverable world model. Experimental results demonstrate that EVOKE significantly improves task performance, generalization to unseen environments, and data efficiency across three backbone architectures, offering new insights into the mechanisms underlying internal knowledge elicitation.
📝 Abstract
Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when predictions are used for planning. However, for LLM agents operating in digital environments, much of this world knowledge is already internalized during pretraining, which shifts the problem from acquiring it to eliciting it. We argue that typical post-training provides little pressure for such elicitation, since supervision under a single goal at each visited state inadvertently drives policies to rely on superficial contextual habits. We introduce EVOKE, a post-training method that supplies this pressure through goal diversity at fixed states. Motivated by theory showing that an agent competent across diverse goals must encode a world model recoverable from its action preferences, EVOKE holds the environment state and interaction history fixed and ranks the same candidate actions under alternative goals, forcing action preferences to change, so that a policy relying on contextual habits or single-goal correlations cannot order them correctly. This implicitly elicits the policy's pretrained world knowledge to inform decisions. We evaluate EVOKE across diverse tasks in three backbones, demonstrating improved task performance, unseen environment generalization, and data efficiency. We further conduct controlled analyses to better understand what drives these gains. These findings offer a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Transferable Decision-Making
World Knowledge
Generalization
Post-training
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Knowledge Elicitation
Goal Diversity
Transferable Decision-Making
LLM Agents
Post-training