Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether large language models (LLMs) align with human judgments in interpreting the emotional valence of news texts, and examines how this alignment varies across sociodemographic groups. Leveraging a representative sample of 3,011 adults from the UK collected via the YouGov platform, the research presents the first systematic evaluation of seven leading LLMs in predicting expressions of sympathy within headlines concerning political and geopolitical conflicts. Results reveal that state-of-the-art models—such as GPT-5.2—exhibit strong overall alignment with human judgments (r = 0.789) and demonstrate robust performance across most subgroups, yet significant disparities in alignment persist between different demographic segments. This work establishes the first large-scale, demographically diverse empirical benchmark for assessing both the affective comprehension capabilities and fairness of LLMs.
📝 Abstract
Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception. Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4 ,Mistral Large 2512). Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants' predispositions regarding the conflict, although there are statistically significant differences between groups. This research, with its robust design and large, demographically diverse dataset, offers the most comprehensive evaluation of LLMs' comprehension of news framing to date. Findings highlight an important, often-ignored aspect of differential alignment: even when aggregate performance is high, AI alignment is not universal -- it may correspond differently with demographic features and cultural norms. Considering or ignoring the need for differential alignment may therefore have significant implications for the development of ethical and useful AI systems.
Problem

Research questions and friction points this paper is trying to address.

AI alignment
emotional perception
sympathetic framing
sociodemographic groups
news framing
Innovation

Methods, ideas, or system contributions that make the work stand out.

differential alignment
emotional framing
large language models
sociodemographic evaluation
AI-human alignment