AI safety needs social scientists

OpenAI Blog News

Summary

OpenAI argues that AI safety research on value alignment requires social scientists to help address how human cognitive biases and inconsistencies affect the data used to train AI systems. The organization proposes human-only experiments as a method to uncover alignment problems before deploying machine learning solutions.

We’ve written a paper arguing that long-term AI safety research needs social scientists to ensure AI alignment algorithms succeed when actual humans are involved. Properly aligning advanced AI systems with human values requires resolving many uncertainties related to the psychology of human rationality, emotion, and biases. The aim of this paper is to spark further collaboration between machine learning and social science researchers, and we plan to hire social scientists to work on this full time at OpenAI.
Original Article
View Cached Full Text

Cached at: 04/20/26, 02:46 PM

# AI safety needs social scientists Source: [https://openai.com/index/ai-safety-needs-social-scientists/](https://openai.com/index/ai-safety-needs-social-scientists/) The goal of long\-term artificial intelligence \(AI\) safety is to ensure that advanced AI systems are aligned with human values—that they reliably do things that people want them to do\. At OpenAI we hope to achieve this by asking people questions about what they want, training machine learning \(ML\) models on this data, and optimizing AI systems to do well according to these learned models\. Examples of this research include[Learning from human preferences⁠\(opens in a new window\)](https://blog.openai.com/deep-reinforcement-learning-from-human-preferences/),[AI safety via debate⁠\(opens in a new window\)](https://blog.openai.com/debate/), and[Learning complex goals with iterated amplification⁠\(opens in a new window\)](https://blog.openai.com/amplifying-ai-training/)\. Unfortunately, human answers to questions about their values may be unreliable\. Humans have limited knowledge and reasoning ability, and exhibit a variety of cognitive biases and ethical beliefs that turn out to be inconsistent on reflection\. We anticipate that different ways of asking questions will interact with human biases in different ways, producing higher or lower quality answers\. For example, judgments about how wrong an action is can vary depending on whether the word “morally” appears in the[question⁠\(opens in a new window\)](http://journal.sjdm.org/10/101109/jdm101109.pdf), and people can make inconsistent choices between gambles if the task they are presented with is[complex⁠\(opens in a new window\)](https://psycnet.apa.org/record/2010-23289-003)\. We have several methods that try to target the reasoning behind human values, including[amplification⁠\(opens in a new window\)](https://blog.openai.com/amplifying-ai-training/)and[debate⁠\(opens in a new window\)](https://blog.openai.com/debate/), but do not know how they behave with real people in realistic situations\. If a problem with an alignment algorithm appears only in natural language discussion of a complex value\-laden question, current ML may be too weak to uncover the issue\. To avoid the limitations of ML, we propose experiments that consist entirely of people, replacing ML agents with people playing the role of those agents\. For example, the[debate⁠\(opens in a new window\)](https://blog.openai.com/debate/)approach to AI alignment involves a game with two AI debaters and a human judge; we can instead use two human debaters and a human judge\. Humans can debate whatever questions we like, and lessons learned in the human case can be transferred to ML\.

Similar Articles

AI safety and alignment

Reddit r/artificial

The article discusses concerns about AI safety and alignment as AI becomes more intelligent and integrated into society, referencing Anthropic's call for a pause to address potential catastrophic risks.

AI safety via debate

OpenAI Blog

OpenAI proposes a novel approach to AI safety where two AI agents debate each other while a human judge evaluates their arguments, allowing humans to supervise AI systems whose behavior is too complex to directly understand. The method leverages debate and adversarial reasoning to align advanced AI with human values and preferences.

Why responsible AI development needs cooperation on safety

OpenAI Blog

OpenAI publishes a policy research paper identifying four strategies to improve industry cooperation on AI safety norms: communicating risks/benefits, technical collaboration, increased transparency, and incentivizing standards. The analysis addresses how competitive pressures could lead to under-investment in safety and proposes mechanisms to align incentives toward safe AI development.

OpenAI safety practices

OpenAI Blog

OpenAI outlines 10 safety practices it actively uses and improves upon, including empirical red-teaming, alignment research, abuse monitoring, and voluntary commitments shared at the AI Seoul Summit. The company emphasizes a balanced, scientific approach to safety integrated into development from the outset.

Our approach to AI safety

OpenAI Blog

OpenAI outlines its comprehensive approach to AI safety, emphasizing rigorous testing, iterative deployment, real-world monitoring, and regulatory engagement to ensure powerful AI systems are built and used safely.