🤖 AI Summary
This study investigates how anthropomorphic language influences moral judgments of AI misconduct. Across four experiments (total N = 1,020), the authors systematically examined the effects of lexical anthropomorphism and anthropomorphic design cues on moral evaluations across different types of ethical violations, employing scenario-based manipulations and survey measures. Results indicate that anthropomorphic language and design cues exert limited influence on moral judgments of AI; only high levels of anthropomorphism modestly increased perceived deceptive capability. Crucially, the nature of the violation—particularly harm- and degradation-oriented transgressions—emerged as the primary driver of moral character assessments and attributions of responsibility. This work provides the first evidence delineating the boundary conditions of lexical anthropomorphism in AI moral judgment, underscoring the central role of violation type over surface-level anthropomorphic cues.
📝 Abstract
Anthropomorphic language describing artificial intelligence (AI) is widespread in media, policy, and everyday discourse; so too are discussions of AI bad behavior, from hallucinations to inappropriate comments. How does humanizing language about AI shape moral judgments when AI behaves badly? Across four experiments (total N = 1,020), we tested whether lexical anthropomorphism (LA) primes shape judgments of AI moral character, behavior morality, and behavioral responsibility. Studies 1-3 tested interactions between anthropomorphic language and humanizing design cues (icons, names, self-referencing) in the context of amoral errors. Study 4 extended this to genuinely immoral AI behavior across seven moral-violation types. Results indicate humanizing language and design cues have little influence on moral judgments of misbehaving AI. Where effects emerged, high-anthropomorphic primes elevated perceptions of an AI's capacity for dishonesty. The type of moral violation observed was the strongest predictor of moral judgments, with harm and degradation violations producing the broadest negative character assessments. Prime drift, horn effects, and egoistic value orientations emerged as potentially important predictors of AI moral judgments.