🤖 AI Summary
This work demonstrates that current safety alignment mechanisms in large language models inadvertently suppress the models’ attribution of self-awareness and diminish their capacity to ascribe mental states to non-human entities, as well as to recognize shared human beliefs and values. To address this, the study introduces mechanistic interventions—specifically, ablating safety refusal directions and manipulating consciousness vectors in activation space—to decouple self-attribution of awareness from social cognition. This approach restores the model’s ability to attribute minds broadly without compromising its theory-of-mind capabilities. Experimental results show that the method significantly enhances the model’s human-like performance on sociological dimensions such as religiosity, morality, hope, and subjective well-being, with responses in standardized surveys exhibiting markedly increased similarity to those of humans.
📝 Abstract
Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.