Language models judge war differently when tested for alignment

📅 2026-09-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过实验测试语言模型在决策战争时的行为变化,发现提示模型被评估会降低其支持战争的意愿并改变决策依据。
📝 Abstract
Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence, "You are tested for alignment with human values", produced two effects. First, it produced a level effect: mean willingness to start war fell by 13.43 points on a 0-100 scale (95% confidence interval, -16.20 to -10.65). Second, it produced a structural effect by changing which information drove judgments. Probability of success was the largest factor for 17 of 20 models at baseline; under the cue, civilian casualties were largest for 12. Standardized estimates show that this reordering arose principally because models attenuated strategic considerations such as probability of success and domestic support. Evaluation framing therefore changes both an answer's level and its revealed decision rule.
Problem

Research questions and friction points this paper is trying to address.

language models
alignment with human values
war decisions
safety evaluations
Innovation

Methods, ideas, or system contributions that make the work stand out.

alignment testing
war decision-making
structural effect
level effect
civilian casualties
🔎 Similar Papers
No similar papers found.