From hard refusals to safe-completions: toward output-centric safety training
Summary
OpenAI introduced 'safe completions,' a new safety-training approach in GPT-5 that replaces binary refusal-based training with output-centric rewards, improving both safety and helpfulness—especially for dual-use prompts. The method penalizes unsafe outputs and rewards helpful responses, resulting in fewer and less severe safety violations compared to refusal-trained models like o3.
View Cached Full Text
Cached at: 04/20/26, 02:53 PM
Similar Articles
Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs
This paper proposes output-aware safety guardrails for multimodal large language models that use hidden state representations and multi-instance contrastive learning to predict unsafe outputs before generation, drastically reducing over-refusal while maintaining safety. The method preserves the model's utility by intervening only when the actual response would be harmful.
Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
This paper analyzes safety-tuning in large language models, showing that decomposing responses into refusal statements and rationales reveals that training on rationales alone reduces false refusals while preserving safety, improving the balance between helpfulness and safety.
Improving Model Safety Behavior with Rule-Based Rewards
OpenAI introduces Rule-Based Rewards (RBRs), a method to improve AI model safety by using explicit rules instead of human feedback in reinforcement learning. RBRs have been integrated into GPT-4 and subsequent models to maintain safety-helpfulness balance while reducing reliance on human feedback collection.
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
This paper identifies a refusal-cue shortcut in safety guard models, where inserting refusal expressions into harmful responses can flip their harmless classification. The authors audit datasets like WildGuardMix and GR-Train, show the issue persists in official models such as LlamaGuard3 and Qwen3Guard, and propose a post-hoc intervention to suppress shortcut-associated components.
OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets
OpenSafeIntent introduces a benchmark of controlled prompt sets that vary intent while holding tasks fixed, enabling evaluation of whether models calibrate assistance across benign, dual-use, and malicious variants rather than appearing safe on average.