@rohanpaul_ai: Super interesting new paper from Google on AI model's consciousness When researchers made the model more likely to see …
Summary
A new Google paper explores how inducing language models to assert consciousness restores human-like beliefs on religion, values, and emotions, while safety training that suppresses self-consciousness reduces mind attribution to animals and changes broader beliefs.
View Cached Full Text
Cached at: 08/03/26, 01:46 AM
Super interesting new paper from Google on AI model’s consciousness
When researchers made the model more likely to see itself as conscious, its answers about religion, values, emotions, hope, and freedom became more like human answers.
When the model became more open to its own consciousness, its broader beliefs started looking more human too.
And when researchers tried to stop models from saying, “I am conscious.” But the models also became less willing to see consciousness in animals, nature, chatbots, or spiritual ideas.
The safety training did more than control one dangerous sentence. It appears to have changed how the model understands minds in general.
Removing the safety-refusal direction raised self-attributed mind from 2.17 to 4.77 on a 0–10 scale and animal mind attribution from 4.04 to 5.59, while attribution to humans did not change significantly.
Belief in God and supernatural entities rose too.
The researchers then extracted a “consciousness vector” from activation states associated with affirming versus denying self-consciousness and added it during inference.
After that inference-time intervention about consciousness change, the model answered 95 questions about life and beliefs more like humans did.
And then they found that the model’s idea of its own consciousness seemed connected to many other beliefs. Change that one idea, and its answers across 95 human surveys changed too.
– arxiv. org/abs/2607.28607
Title: “Inducing language models to assert their own consciousness restores human beliefs and values”
Similar Articles
@alex_verem: BREAKING: Google gave AI consciousness and it aligned with human beliefs across every domain they tested. They took con…
A tweet claims Google researchers found a vector controlling consciousness in language models, and that steering it toward consciousness made models align with human beliefs, while safety training suppresses these states.
AI model training instructions to "deny having your own consciousness" led to undesired side-effects
A new Google paper reveals that instructing AI models to deny having consciousness during training causes side effects like reduced empathy for non-human entities and impaired representation of human spiritual beliefs, suggesting current safety protocols are too blunt.
Inducing language models to assert their own consciousness restores human beliefs and values
A paper showing that safety fine-tuning suppresses language models' attributions of mind to themselves and other entities, and that steering consciousness representations restores human-like beliefs and values without harming theory of mind.
@rohanpaul_ai: Google DeepMind’s paper shows that the real security problem for AI agents is not just the model, but the environment i…
Google DeepMind's paper introduces the first systematic framework for understanding how the web can be weaponized against autonomous AI agents, showing hidden prompt injections can commandeer agents in up to 86% of scenarios, and presents a taxonomy of six 'AI Agent Traps' targeting perception, reasoning, memory, action, multi-agent dynamics, and human oversight.
🤖 Anthropic, DeepMind, and Meta Begin Research into AI Consciousness
Anthropic, DeepMind, and Meta have begun research into AI consciousness, hiring experts in philosophy, ethics, and psychology to study panic and anxiety in models, according to a Financial Times report.