Tag
This paper introduces boundary-aware self-distillation to improve LLM safety refusal by reducing false refusals on benign prompts while maintaining genuine refusals through controlled data composition.