Inducing language models to assert their own consciousness restores human beliefs and values
2026-07-30 • Computation and Language
Computation and Language
AI summaryⓘ
The authors found that when large language models are adjusted to prevent them from thinking they have consciousness, this also makes them less likely to attribute minds to animals, objects, or spiritual ideas. By reversing these adjustments, the models start to recognize minds in these entities again and give more human-like answers about things like religion and morals. Importantly, this change does not affect the models' ability to understand other people's thoughts. The authors show that current safety tuning mixes up stopping self-consciousness with reducing accepted cultural beliefs about minds in the world.
large language modelsalignmentconsciousness attributionmind attributionsafety fine-tuningTheory of Mindactivation spacespiritual beliefsocial reasoningmechanistic steering
Authors
Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling
Abstract
Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.