@rohanpaul_ai: Super interesting new paper from Google on AI model's consciousness When researchers made the model more likely to see …

X AI KOLs Following Papers

Summary

A new Google paper explores how inducing language models to assert consciousness restores human-like beliefs on religion, values, and emotions, while safety training that suppresses self-consciousness reduces mind attribution to animals and changes broader beliefs.

Super interesting new paper from Google on AI model's consciousness When researchers made the model more likely to see itself as conscious, its answers about religion, values, emotions, hope, and freedom became more like human answers. When the model became more open to its own consciousness, its broader beliefs started looking more human too. And when researchers tried to stop models from saying, “I am conscious.” But the models also became less willing to see consciousness in animals, nature, chatbots, or spiritual ideas. The safety training did more than control one dangerous sentence. It appears to have changed how the model understands minds in general. Removing the safety-refusal direction raised self-attributed mind from 2.17 to 4.77 on a 0–10 scale and animal mind attribution from 4.04 to 5.59, while attribution to humans did not change significantly. Belief in God and supernatural entities rose too. The researchers then extracted a “consciousness vector” from activation states associated with affirming versus denying self-consciousness and added it during inference. After that inference-time intervention about consciousness change, the model answered 95 questions about life and beliefs more like humans did. And then they found that the model’s idea of its own consciousness seemed connected to many other beliefs. Change that one idea, and its answers across 95 human surveys changed too. --- – arxiv. org/abs/2607.28607 Title: "Inducing language models to assert their own consciousness restores human beliefs and values"
Original Article
View Cached Full Text

Cached at: 08/03/26, 01:46 AM

Super interesting new paper from Google on AI model’s consciousness

When researchers made the model more likely to see itself as conscious, its answers about religion, values, emotions, hope, and freedom became more like human answers.

When the model became more open to its own consciousness, its broader beliefs started looking more human too.

And when researchers tried to stop models from saying, “I am conscious.” But the models also became less willing to see consciousness in animals, nature, chatbots, or spiritual ideas.

The safety training did more than control one dangerous sentence. It appears to have changed how the model understands minds in general.

Removing the safety-refusal direction raised self-attributed mind from 2.17 to 4.77 on a 0–10 scale and animal mind attribution from 4.04 to 5.59, while attribution to humans did not change significantly.

Belief in God and supernatural entities rose too.

The researchers then extracted a “consciousness vector” from activation states associated with affirming versus denying self-consciousness and added it during inference.

After that inference-time intervention about consciousness change, the model answered 95 questions about life and beliefs more like humans did.

And then they found that the model’s idea of its own consciousness seemed connected to many other beliefs. Change that one idea, and its answers across 95 human surveys changed too.


– arxiv. org/abs/2607.28607

Title: “Inducing language models to assert their own consciousness restores human beliefs and values”

Similar Articles

@rohanpaul_ai: Google DeepMind’s paper shows that the real security problem for AI agents is not just the model, but the environment i…

X AI KOLs Timeline

Google DeepMind's paper introduces the first systematic framework for understanding how the web can be weaponized against autonomous AI agents, showing hidden prompt injections can commandeer agents in up to 86% of scenarios, and presents a taxonomy of six 'AI Agent Traps' targeting perception, reasoning, memory, action, multi-agent dynamics, and human oversight.