Training Language Models To Be Coherent Decision-Makers

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of belief instability and the inability to recognize missing information that language models frequently encounter during decision-making. To tackle these issues, this work proposes a routing architecture that decouples decision completeness discrimination from utility optimization. Through supervised fine-tuning and probability elicitation techniques, the model is trained to maintain consistent beliefs while maximizing expected utility. The proposed approach significantly enhances decision coherence and the capacity to identify missing information. Furthermore, it achieves effective cross-domain generalization in unseen domains, establishing a novel paradigm for reliable decision-making under incomplete information.
📝 Abstract
Reliable decision-making requires more than accurate prediction: a model must preserve its beliefs, apply the relevant utilities, and recognize when the information needed to justify an action is missing. We study whether language models can learn this decision procedure from supervised fine-tuning and generalize it across domains and differing natural-language expressions of the decision challenge. Across 20 datasets, we explore challenges of belief instability and decision-making errors by first eliciting probabilities of outcomes and then varying only the utilities and the framing of the decision problems, while holding the evidence fixed. We train models to preserve elicited beliefs while selecting the action that maximizes expected utility, and evaluate transfer to unseen application domains, held-out framings, and different classes of payoff structures. We further introduce incomplete-information settings in which required utilities are withheld and replaced with irrelevant text, testing whether models can distinguish missing decision-relevant information from merely additional context. We find that targeted fine-tuning substantially improves coherent decision-making and that in many situations, learning transfers across domains and framings to situations unobserved during training. Further, models trained for decidability learn to identify when action cannot be justified based on missing information. Finally, we show the value of a routed system that considers separately the recognition of decision completeness and utility-sensitive decision execution.
Problem

Research questions and friction points this paper is trying to address.

Language Models
Coherent Decision-Making
Belief Stability
Expected Utility
Incomplete Information
Innovation

Methods, ideas, or system contributions that make the work stand out.

coherent decision-making
supervised fine-tuning
expected utility maximization
incomplete information
routed system