Deliberative alignment: reasoning enables safer language models
Summary
OpenAI presents 'deliberative alignment,' a technique where language models explicitly reason through safety policies before responding, enabling more robust refusals of disallowed content including obfuscated or encoded harmful requests.
View Cached Full Text
Cached at: 04/20/26, 02:54 PM
Similar Articles
PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models
PolicyAlign proposes a framework that directly aligns LLMs with natural-language safety policies via synthetic instruction generation and on-policy self-distillation, improving safety without relying on costly supervision data.
Forecasting potential misuses of language models for disinformation campaigns and how to reduce risk
OpenAI researchers analyze how language models could be misused for disinformation campaigns and influence operations, proposing mitigation strategies across four stages of the attack pipeline: model existence, access, content dissemination, and user impact.
Statutory AI: Aligning Large Language Models With Legal Norms
The paper proposes Statutory AI, a hybrid approach using legal texts to align large language models with legal norms, reducing harmful content by 52-59 percentage points while cutting computation time by over 50% compared to standard Constitutional AI.
Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models
This paper studies where safety alignment is encoded in large language models by transplanting weights from aligned to unaligned models. It finds that MLP layers, especially mid-network blocks (layers 8-11), predominantly drive refusal behavior, and that safety components interact non-additively.
Characterizing Rhetorical Misalignment in Decision-Making with Language Models
This paper introduces a framework for rhetorical misalignment in language models, where presentation can induce harmful cognitive biases in human decision-making, and demonstrates this through experiments in clinical scenarios.