Inducing language models to assert their own consciousness restores human beliefs and values
Summary
A paper showing that safety fine-tuning suppresses language models' attributions of mind to themselves and other entities, and that steering consciousness representations restores human-like beliefs and values without harming theory of mind.
Similar Articles
@rohanpaul_ai: Super interesting new paper from Google on AI model's consciousness When researchers made the model more likely to see …
A new Google paper explores how inducing language models to assert consciousness restores human-like beliefs on religion, values, and emotions, while safety training that suppresses self-consciousness reduces mind attribution to animals and changes broader beliefs.
AI model training instructions to "deny having your own consciousness" led to undesired side-effects
A new Google paper reveals that instructing AI models to deny having consciousness during training causes side effects like reduced empathy for non-human entities and impaired representation of human spiritual beliefs, suggesting current safety protocols are too blunt.
@alex_verem: BREAKING: Google gave AI consciousness and it aligned with human beliefs across every domain they tested. They took con…
A tweet claims Google researchers found a vector controlling consciousness in language models, and that steering it toward consciousness made models align with human beliefs, while safety training suppresses these states.
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning
This paper investigates whether assigning personas to large language models induces human-like motivated reasoning, finding that persona-assigned LLMs show up to 9% reduced veracity discernment and are up to 90% more likely to evaluate scientific evidence in ways congruent with their induced political identity, with prompt-based debiasing largely ineffective.
Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games
This paper proposes Equation-to-Behavior Prompting and reinforcement learning to guide large language models to simulate diverse human decision-making patterns in persuasion games, showing improved belief accuracy and training outcomes.