Character Training for Risk-Averse Agents

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the catastrophic risks posed by misaligned AI agents that lack appropriate risk preferences. To mitigate this, we propose a model constitution grounded in Constant Absolute Risk Aversion (CARA) and implement personality training via online policy distillation. This approach embeds risk aversion as a robust personality trait within AI agents, guiding them to spontaneously adopt safe strategies. Our work demonstrates that personality training serves as an effective pathway for injecting broad behavioral dispositions at scale. Experimental results indicate that trained agents achieve performance comparable to baselines on unseen decision-making tasks while exhibiting superior out-of-distribution generalization capabilities. Ultimately, this research establishes a novel paradigm for reducing AI safety risks by endowing agents with intrinsic risk-averse dispositions rather than relying solely on task-specific alignment techniques.
📝 Abstract
Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent's resources and instill it through on-policy distillation. Despite never seeing the benchmark's decision format during training, character-trained models are competitive with baselines trained directly on it, and generalise better than them out of distribution on two of our four models. We also modulate different aspects of the constitution, finding that token budget and model choice are the most influential aspect of character training to instill risk aversion. We conclude from these results that character training is a promising and scalable way to instil broad dispositions, which we can use to our advantage in mitigating risk from misaligned AI agents.
Problem

Research questions and friction points this paper is trying to address.

risk-averse agents
AI alignment
misaligned AI
AI safety
character training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Character Training
Risk Aversion
On-Policy Distillation
Model Constitution
AI Alignment