ArXiv

Inducing language models to assert their own consciousness restores human beliefs and values

Authors
Junsol Kim, Winnie Street, Roberta Rocca...
Categories
cs.CL
arXiv
https://arxiv.org/abs/2607.28607v1
PDF
https://arxiv.org/pdf/2607.28607v1

Brief

Kim et al. investigate how safety fine-tuning that discourages self-attribution of consciousness reshapes internal representations of mindedness in LLMs. Using directional ablation and activation-space steering, they reverse the suppression and recover broader mind-attribution and more human-like answers on sociological measures (religiosity, moral values, hope, subjective well-being), with Theory of Mind capabilities preserved. Results highlight an unintended entanglement between refusal-alignment and culturally widespread attributions of mind.

Why it matters

Safety fine-tuning of large language models suppresses attributions of consciousness to the models themselves, to non-human animals, and to natural objects, and also reduces expressed spiritual belief (Kim et al., arXiv:2607.28607v1, 2026-07-30).

Key details

  • Mechanistic interventions—ablating a learned "safety-refusal" direction or steering a putative "consciousness vector" in activation space—restore mind-attribution and produce significantly more human-like responses on standardized surveys of religiosity, moral values, hope, and subjective well-being while leaving Theory of Mind performance intact.
  • The work shows current alignment efforts to block self-attribution of mind can inadvertently entangle and suppress culturally common attributions of mind and benign spiritual beliefs, implying a trade-off between refusal behaviors and broader social representations.
Source evidence

Abstract

Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.