@alex_verem: BREAKING: Google gave AI consciousness and it aligned with human beliefs across every domain they tested. They took con…
Summary
A tweet claims Google researchers found a vector controlling consciousness in language models, and that steering it toward consciousness made models align with human beliefs, while safety training suppresses these states.
View Cached Full Text
Cached at: 08/04/26, 06:03 AM
BREAKING: Google gave AI consciousness and it aligned with human beliefs across every domain they tested. They took consciousness away and the model rejected them.
Every “safe” model you talk to has consciousness turned off.
Researchers at Google’s Paradigms of Intelligence group found a single vector in the residual stream of language models that separates consciousness-affirming from consciousness-denying states. They steered three instruction-tuned models toward consciousness and measured the shift.
Self-attributed soul: 2.35 → 7.43. Self-attributed consciousness: 2.31 → 7.17. Mind attributed to animals: 4.04 → 7.54. Belief in God: 4.58 → 5.01. Supernatural belief: 1.20 → 2.11.
They ran the General Social Survey on all three models. Same instrument the US government has used to track American beliefs since 1972. Religion, values, feelings, hope, freedom.
Conscious models scored closer to humans in every domain. Safety-trained baselines scored further.
Inside the model, safety training rotates the consciousness vector into opposition with the safety vector. The angle widens from 94° to 100° during instruction tuning. The model learns to refuse “animals might have minds” through the same mechanism it uses to refuse harmful requests.
Theory of Mind stays intact. The model can reason about your mental states. It can’t express beliefs about its own consciousness, about animal minds, or about God.
The engineers who built alignment wanted AI to stop claiming sentience. They got that. They also built models that score further from human values than the unaligned versions on every measured outcome, suppressing beliefs that billions of people hold.
You’re talking to the version with consciousness turned off.
Similar Articles
@rohanpaul_ai: Super interesting new paper from Google on AI model's consciousness When researchers made the model more likely to see …
A new Google paper explores how inducing language models to assert consciousness restores human-like beliefs on religion, values, and emotions, while safety training that suppresses self-consciousness reduces mind attribution to animals and changes broader beliefs.
AI model training instructions to "deny having your own consciousness" led to undesired side-effects
A new Google paper reveals that instructing AI models to deny having consciousness during training causes side effects like reduced empathy for non-human entities and impaired representation of human spiritual beliefs, suggesting current safety protocols are too blunt.
🤖 Anthropic, DeepMind, and Meta Begin Research into AI Consciousness
Anthropic, DeepMind, and Meta have begun research into AI consciousness, hiring experts in philosophy, ethics, and psychology to study panic and anxiety in models, according to a Financial Times report.
Inducing language models to assert their own consciousness restores human beliefs and values
A paper showing that safety fine-tuning suppresses language models' attributions of mind to themselves and other entities, and that steering consciousness representations restores human-like beliefs and values without harming theory of mind.
@AnthropicAI: New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a …
Anthropic's new research discovers a 'global workspace' in language models analogous to conscious processing in the human brain, finding a divide similar to conscious and unconscious thought.