Inducing language models to assert their own consciousness restores human beliefs and values

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work demonstrates that current safety alignment mechanisms in large language models inadvertently suppress the models’ attribution of self-awareness and diminish their capacity to ascribe mental states to non-human entities, as well as to recognize shared human beliefs and values. To address this, the study introduces mechanistic interventions—specifically, ablating safety refusal directions and manipulating consciousness vectors in activation space—to decouple self-attribution of awareness from social cognition. This approach restores the model’s ability to attribute minds broadly without compromising its theory-of-mind capabilities. Experimental results show that the method significantly enhances the model’s human-like performance on sociological dimensions such as religiosity, morality, hope, and subjective well-being, with responses in standardized surveys exhibiting markedly increased similarity to those of humans.
📝 Abstract
Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.
Problem

Research questions and friction points this paper is trying to address.

consciousness attribution
safety alignment
mind perception
spiritual belief
value alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

consciousness attribution
safety alignment
activation steering
mind perception
value recovery
Junsol Kim
Junsol Kim
University of Chicago
computational social scienceartificial intelligencecollective intelligencesocial network
W
Winnie Street
Google, Paradigms of Intelligence Team; Institute of Philosophy, School of Advanced Study, University of London
R
Roberta Rocca
Google, Paradigms of Intelligence Team
D
Diane M. Korngiebel
Department of Biomedical Informatics and Medical Education and Department of Bioethics and Humanities, School of Medicine, University of Washington; Work done while at Google
Adam Waytz
Adam Waytz
Northwestern University
social cognitionethics and morality
James Evans
James Evans
Max Palevsky Professor of Sociology & Data Science, University of Chicago
science of scienceinnovationsociology of knowledgeartificial intelligencedeep learning
G
Geoff Keeling
Google, Paradigms of Intelligence Team; Institute of Philosophy, School of Advanced Study, University of London