Tag
The article discusses the limitations of topic-level safety guards in AI models and introduces a new paper proposing boundary-aware self-distillation for controlled LLM safety refusal, focusing on refusing specific harmful subsets within topics rather than entire topics.